Fetching the paper…
Reading the bibliography…
In the field of 3D object detection tasks, fusing heterogeneous features from LiDAR and camera sensors into a unified Bird's Eye View (BEV) representation is a widely adopted paradigm.
R. Hadsell, S. Chopra, and Y. LeCun, “Dimensionality reduction by learning an invariant mapping,” in 2006 IEEE computer society conference on computer vision and pattern recognition (CVPR’06) , vol. 2. IEEE, 2006, pp. 1735–1742
2006
Earlier work this paper cites.
2014
Earlier work this paper cites.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 2980–2988
2017
Earlier work this paper cites.
C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “Pointnet: Deep learning on point sets for 3D classification and segmentation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 652–660
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 2961–2969
2017
Earlier work this paper cites.
Y. Yan, Y. Mao, and B. Li, “Second: Sparsely embedded convolutional detection,” Sensors , vol. 18, no. 10, p. 3337, 2018
2018
Earlier work this paper cites.
Y. Zhou and O. Tuzel, “Voxelnet: End-to-end learning for point cloud based 3D object detection,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 4490–4499
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. H. Lang, S. Vora, H. Caesar, L. Zhou, J. Yang, and O. Beijbom, “Pointpillars: Fast encoders for object detection from point clouds,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 12 697–12 705
2019
Earlier work this paper cites.
S. Shi, X. Wang, and H. Li, “Pointrcnn: 3D object proposal generation and detection from point cloud,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 770–779
2019
Earlier work this paper cites.
Z. Liu, H. Tang, Y. Lin, and S. Han, “Point-voxel cnn for efficient 3D deep learning,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
V. A. Sindagi, Y. Zhou, and O. Tuzel, “Mvx-net: Multimodal voxelnet for 3D object detection,” in 2019 International Conference on Robotics and Automation (ICRA) . IEEE, 2019, pp. 7276–7282
2019
Earlier work this paper cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al. , “Pytorch: An imperative style, high-performance deep learning library,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Philion and S. Fidler, “Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3D,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XIV 16 . Springer, 2020, pp. 194–210
2020
Earlier work this paper cites.
F. Wang, Y. Zhuang, H. Zhang, and H. Gu, “Real-time 3-d semantic scene parsing with lidar sensors,” IEEE Transactions on Cybernetics , vol. 52, no. 3, pp. 1351–1363, 2020
2020
Earlier work this paper cites.
S. Vora, A. H. Lang, B. Helou, and O. Beijbom, “Pointpainting: Sequential fusion for 3D object detection,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 4604–4612
2020
Earlier work this paper cites.
T. Huang, Z. Liu, X. Chen, and X. Bai, “Epnet: Enhancing point features with image semantics for 3D object detection,” in European Conference on Computer Vision . Springer, 2020, pp. 35–52
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom, “NuScenes: A multimodal dataset for autonomous driving,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 11 621–11 631
2020
Earlier work this paper cites.
A. Jaiswal, A. R. Babu, M. Z. Zadeh, D. Banerjee, and F. Makedon, “A survey on contrastive self-supervised learning,” Technologies , vol. 9, no. 1, p. 2, 2020
2020
Earlier work this paper cites.
Y. Tian, C. Sun, B. Poole, D. Krishnan, C. Schmid, and P. Isola, “What makes for good views for contrastive learning?” Advances in neural information processing systems , vol. 33, pp. 6827–6839, 2020
2020
Earlier work this paper cites.
P. H. Le-Khac, G. Healy, and A. F. Smeaton, “Contrastive representation learning: A framework and review,” Ieee Access , vol. 8, pp. 193 907–193 934, 2020
2020
Earlier work this paper cites.
O. D. Team, “OpenPCDet: An Open-source Toolbox for 3D Object Detection from Point Clouds,” https://github.com/open-mmlab/OpenPCDet , 2020
2020
Earlier work this paper cites.
D. Feng, S. Han, H. Xu, X. Liang, and X. Tan, “Point-guided contrastive learning for monocular 3-d object detection,” IEEE Transactions on Cybernetics , vol. 53, no. 2, pp. 954–966, 2021
2021
Earlier work this paper cites.
J. Deng, S. Shi, P. Li, W. Zhou, Y. Zhang, and H. Li, “Voxel r-cnn: Towards high performance voxel-based 3D object detection,” in Proceedings of the AAAI conference on artificial intelligence , vol. 35, no. 2, 2021, pp. 1201–1209
2021
Earlier work this paper cites.
C. Wang, C. Ma, M. Zhu, and X. Yang, “Pointaugmenting: Cross-modal augmentation for 3D object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 11 794–11 803
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
T. Yin, X. Zhou, and P. Krähenbühl, “Multimodal virtual point 3D detection,” Advances in Neural Information Processing Systems , vol. 34, pp. 16 494–16 507, 2021
2021
Cited alongside, same era.
T. Xiao, C. J. Reed, X. Wang, K. Keutzer, and T. Darrell, “Region similarity representation learning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 10 539–10 548
2021
Cited alongside, same era.
X. Wang, R. Zhang, C. Shen, T. Kong, and L. Li, “Dense contrastive learning for self-supervised visual pre-training,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 3024–3033
2021
Cited alongside, same era.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 10 012–10 022
2021
Cited alongside, same era.
Z. Song, C. Jia, L. Yang, H. Wei, and L. Liu, “Graphalign++: An accurate feature alignment by graph matching for multi-modal 3D object detection,” IEEE Transactions on Circuits and Systems for Video Technology , 2023
2023
Later among the works it cites.
Z. Song, G. Zhang, J. Xie, L. Liu, C. Jia, S. Xu, and Z. Wang, “VoxelNextFusion: A Simple, Unified, and Effective Voxel Fusion Framework for Multimodal 3-D Object Detection,” IEEE Transactions on Geoscience and Remote Sensing , vol. 61, pp. 1–12, 2023
2023
Later among the works it cites.
Y. Li, Z. Ge, G. Yu, J. Yang, Z. Wang, Y. Shi, J. Sun, and Z. Li, “Bevdepth: Acquisition of reliable depth for multi-view 3D object detection,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, no. 2, 2023, pp. 1477–1485
2023
Later among the works it cites.
Q. Cai, Y. Pan, T. Yao, C.-W. Ngo, and T. Mei, “ObjectFusion: Multi-modal 3D Object Detection with Object-Centric Fusion,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , October 2023, pp. 18 067–18 076
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Yin, X. Zhou, and P. Krahenbuhl, “Center-based 3D Object Detection and Tracking,” in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , Jun 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
B. Wilson, W. Qi, T. Agarwal, J. Lambert, J. Singh, S. Khandelwal, B. Pan, R. Kumar, A. Hartnett, J. K. Pontes, D. Ramanan, P. Carr, and J. Hays, “Argoverse 2: Next Generation Datasets for Self-Driving Perception and Forecasting,” in Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks (NeurIPS Datasets and Benchmarks 2021) , 2021
2021
Cited alongside, same era.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin Transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 10 012–10 022
2021
Cited alongside, same era.
T. Liang, H. Xie, K. Yu, Z. Xia, Z. Lin, Y. Wang, T. Tang, B. Wang, and Z. Tang, “Bevfusion: A simple and robust LiDAR-camera fusion framework,” Advances in Neural Information Processing Systems , vol. 35, pp. 10 421–10 434, 2022
2022
Cited alongside, same era.
K. Liu, Z. Gao, F. Lin, and B. M. Chen, “Fg-net: A fast and accurate framework for large-scale lidar point cloud understanding,” IEEE Transactions on Cybernetics , vol. 53, no. 1, pp. 553–564, 2022
2022
Cited alongside, same era.
X. Bai, Z. Hu, X. Zhu, Q. Huang, Y. Chen, H. Fu, and C.-L. Tai, “Transfusion: Robust LiDAR-camera fusion for 3D object detection with transformers,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 1090–1099
2022
Cited alongside, same era.
Y. Li, A. W. Yu, T. Meng, B. Caine, J. Ngiam, D. Peng, J. Shen, Y. Lu, D. Zhou, Q. V. Le et al. , “Deepfusion: LiDAR-camera deep fusion for multi-modal 3D object detection,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 17 182–17 191
2022
Cited alongside, same era.
2023
Later among the works it cites.
F. Zhan, Y. Yu, R. Wu, J. Zhang, S. Lu, L. Liu, A. Kortylewski, C. Theobalt, and E. Xing, “Multimodal image synthesis and editing: A survey and taxonomy,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2023
2023
Later among the works it cites.
P. Xu, X. Zhu, and D. A. Clifton, “Multimodal learning with transformers: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2023
2023
Later among the works it cites.
M. U. Khattak, H. Rasheed, M. Maaz, S. Khan, and F. S. Khan, “Maple: Multi-modal prompt learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 19 113–19 122
2023
Later among the works it cites.
Y. Dong, C. Kang, J. Zhang, Z. Zhu, Y. Wang, X. Yang, H. Su, X. Wei, and J. Zhu, “Benchmarking robustness of 3D object detection to common corruptions,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 1022–1032
2023
Later among the works it cites.
R. Fan, M. Poggi, and S. Mattoccia, “Contrastive Learning for Depth Prediction,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 3225–3236
2023
Later among the works it cites.
Y. Chen, J. Liu, X. Zhang, X. Qi, and J. Jia, “VoxelNeXt: Fully Sparse VoxelNet for 3D Object Detection and Tracking,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023
2023
Later among the works it cites.
X. Chen, T. Zhang, Y. Wang, Y. Wang, and H. Zhao, “FUTR3D: A unified sensor fusion framework for 3D detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 172–181
2023
Later among the works it cites.
J. Yan, Y. Liu, J. Sun, F. Jia, S. Li, T. Wang, and X. Zhang, “Cross modal transformer: Towards fast and robust 3d object detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 18 268–18 278
2023
Later among the works it cites.
H. Wang, H. Tang, S. Shi, A. Li, Z. Li, B. Schiele, and L. Wang, “UniTR: A unified and efficient multi-modal transformer for bird’s-eye-view representation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 6792–6802
2023
Later among the works it cites.
Y. Chen, Z. Yu, Y. Chen, S. Lan, A. Anandkumar, J. Jia, and J. M. Alvarez, “FocalFormer3D: focusing on hard instance for 3d object detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 8394–8405
2023
Later among the works it cites.
Y. Jiao, Z. Jie, S. Chen, J. Chen, L. Ma, and Y.-G. Jiang, “Msmdfusion: Fusing lidar and camera at multiple scales with multi-depth seeds for 3d object detection,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2023, pp. 21 643–21 652
2023
Later among the works it cites.
Y. Xie, C. Xu, M.-J. Rakotosaona, P. Rim, F. Tombari, K. Keutzer, M. Tomizuka, and W. Zhan, “Sparsefusion: Fusing multi-modal sparse representations for multi-sensor 3d object detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 17 591–17 602
2023
Later among the works it cites.
2024
Closest in time.
2024
Closest in time.
Y. Chen, G. Cai, Z. Song, Z. Liu, B. Zeng, J. Li, and Z. Wang, “Lvp: Leverage virtual points in multi-modal early fusion for 3d object detection,” IEEE Transactions on Geoscience and Remote Sensing , 2024
2024
Closest in time.
Z. Yang, Y. Zhou, L. Xie, and J. Yang, “T3dnet: Compressing point cloud models for lightweight 3-d recognition,” IEEE Transactions on Cybernetics , 2024
2024
Closest in time.
S. Xu, S. Jiang, F. Li, L. Liu, Z. Song, B. Yang, and Z.-x. Yang, “Sparseinteraction: Sparse semantic guidance for radar and camera 3d object detection,” in Proceedings of the 32nd ACM International Conference on Multimedia , 2024, pp. 9224–9233
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
J. Bi, H. Wei, G. Zhang, K. Yang, and Z. Song, “DyFusion: Cross-attention 3d object detection with dynamic fusion,” IEEE Latin America Transactions , vol. 22, no. 2, pp. 106–112, 2024
2024
Closest in time.
2024
Closest in time.
N. Carlini, M. Nasr, C. A. Choquette-Choo, M. Jagielski, I. Gao, P. W. W. Koh, D. Ippolito, F. Tramer, and L. Schmidt, “Are aligned neural networks adversarially aligned?” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
S. Xu, F. Li, Z. Song, J. Fang, S. Wang, and Z.-X. Yang, “Multi-Sem fusion: multimodal semantic fusion for 3D object detection,” IEEE Transactions on Geoscience and Remote Sensing , 2024
2024
Closest in time.
2024
Closest in time.
J. Zhou, Q. Zhu, Y. Wang, M. Feng, J. Liu, J. Huang, and A. Mian, “A state space model for multiobject full 3-d information estimation from rgb-d images,” IEEE Transactions on Cybernetics , 2025
2025
Closest in time.