Fetching the paper…
Reading the bibliography…
Existing top-performance autonomous driving systems typically rely on the multi-modal fusion strategy for reliable scene understanding.
M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman, “The pascal visual object classes (voc) challenge,” IJCV , 2010
2010
Earlier work this paper cites.
T. Y. Lin, M. Maire, S. Belongie, J. Hays, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in ECCV , 2014
2014
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS , 2017
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in ICCV , 2017
2017
Earlier work this paper cites.
Y. Zhou and O. Tuzel, “Voxelnet: End-to-end learning for point cloud based 3d object detection,” in CVPR , 2018
2018
Earlier work this paper cites.
Y. Yan, Y. Mao, and B. Li, “Second: Sparsely embedded convolutional detection,” Sensors , 2018
2018
Earlier work this paper cites.
J. Ku, M. Mozifian, J. Lee, A. Harakeh, and S. L. Waslander, “Joint 3d proposal generation and object detection from view aggregation,” in IROS , 2018
2018
Earlier work this paper cites.
J. Ku, A. Harakeh, and S. L. Waslander, “In defense of classical image processing: Fast depth completion on the cpu,” in CRV , 2018
2018
Earlier work this paper cites.
Y. Wang, W.-L. Chao, D. Garg, B. Hariharan, M. Campbell, and K. Weinberger, “Pseudo-lidar from visual depth estimation: Bridging the gap in 3d object detection for autonomous driving,” in CVPR , 2019
2019
Earlier work this paper cites.
A. H. Lang, S. Vora, H. Caesar, L. Zhou, J. Yang, and O. Beijbom, “Pointpillars: Fast encoders for object detection from point clouds,” in CVPR , 2019
2019
Earlier work this paper cites.
P. Ramachandran, N. Parmar, A. Vaswani, I. Bello, A. Levskaya, and J. Shlens, “Stand-alone self-attention in vision models,” in NeurIPS , 2019
2019
Earlier work this paper cites.
Z. Cai and N. Vasconcelos, “Cascade r-cnn: high quality object detection and instance segmentation,” IEEE TPAMI , 2019
2019
Earlier work this paper cites.
B. Zhu, Z. Jiang, X. Zhou, Z. Li, and G. Yu, “Class-balanced grouping and sampling for point cloud 3d object detection,” arXiv preprint , 2019
2019
Earlier work this paper cites.
W. Zeng, W. Luo, S. Suo, A. Sadat, B. Yang, S. Casas, and R. Urtasun, “End-to-end interpretable neural motion planner,” in CVPR , 2019
2019
Earlier work this paper cites.
S. Vora, A. H. Lang, B. Helou, and O. Beijbom, “Pointpainting: Sequential fusion for 3d object detection,” in CVPR , 2020
2020
Earlier work this paper cites.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in ECCV , 2020
2020
Earlier work this paper cites.
A. Bewley, P. Sun, T. Mensink, D. Anguelov, and C. Sminchisescu, “Range conditioned dilated convolutions for scale invariant 3d object detection,” arXiv preprint , 2020
2020
Earlier work this paper cites.
T. Huang, Z. Liu, X. Chen, and X. Bai, “Epnet: Enhancing point features with image semantics for 3d object detection,” in ECCV , 2020
2020
Earlier work this paper cites.
C. R. Qi, X. Chen, O. Litany, and L. J. Guibas, “Imvotenet: Boosting 3d object detection in point clouds with image votes,” in CVPR , 2020
2020
Earlier work this paper cites.
J. Philion and S. Fidler, “Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d,” in ECCV , 2020
2020
Earlier work this paper cites.
H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom, “nuscenes: A multimodal dataset for autonomous driving,” in CVPR , 2020
2020
Earlier work this paper cites.
M. Contributors, “MMDetection3D: OpenMMLab next-generation platform for general 3D object detection,” https://github.com/open-mmlab/mmdetection3d
2020
Earlier work this paper cites.
P. Sun, H. Kretzschmar, X. Dotiwalla, A. Chouard, V. Patnaik, P. Tsui, J. Guo, Y. Zhou, Y. Chai, B. Caine et al. , “Scalability in perception for autonomous driving: Waymo open dataset,” in CVPR , 2020
2020
Earlier work this paper cites.
C. Wang, C. Ma, M. Zhu, and X. Yang, “Pointaugmenting: Cross-modal augmentation for 3d object detection,” in CVPR , 2021
2021
Earlier work this paper cites.
T. Yin, X. Zhou, and P. Krähenbühl, “Multimodal virtual point 3d detection,” NeurIPS , 2021
2021
Cited alongside, same era.
S. Xu, D. Zhou, J. Fang, J. Yin, B. Zhou, and L. Zhang, “FusionPainting: Multimodal fusion with adaptive attention for 3d object detection,” ITSC , 2021
2021
Cited alongside, same era.
J. Huang, G. Huang, Z. Zhu, and D. Du, “Bevdet: High-performance multi-camera 3d object detection in bird-eye-view,” arXiv preprint , 2021
2021
Cited alongside, same era.
C. Reading, A. Harakeh, J. Chae, and S. L. Waslander, “Categorical depth distributionnetwork for monocular 3d object detection,” in CVPR , 2021
2021
Cited alongside, same era.
I. Misra, R. Girdhar, and A. Joulin, “An End-to-End Transformer Model for 3D Object Detection,” in ICCV , 2021
2021
Cited alongside, same era.
Y. Jiang, L. Zhang, Z. Miao, X. Zhu, J. Gao, W. Hu, and Y.-G. Jiang, “Polarformer: Multi-camera 3d object detection with polar transformers,” arXiv preprint , 2022
2022
Later among the works it cites.
X. Bai, Z. Hu, X. Zhu, Q. Huang, Y. Chen, H. Fu, and C.-L. Tai, “Transfusion: Robust lidar-camera fusion for 3d object detection with transformers,” in CVPR , 2022
2022
Later among the works it cites.
X. Chen, T. Zhang, Y. Wang, Y. Wang, and H. Zhao, “Futr3d: A unified sensor fusion framework for 3d detection,” arXiv preprint , 2022
2022
Later among the works it cites.
S. Hu, L. Chen, P. Wu, H. Li, J. Yan, and D. Tao, “St-p3: End-to-end vision-based autonomous driving via spatial-temporal feature learning,” in ECCV , 2022
2022
Later among the works it cites.
H. Liu, T. Lu, Y. Xu, J. Liu, W. Li, and L. Chen, “Camliflow: bidirectional camera-lidar fusion for joint optical flow and scene flow estimation,” in CVPR , 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Fan, X. Xiong, F. Wang, N. Wang, and Z. Zhang, “Rangedet: In defense of range view for lidar-based 3d object detection,” in ICCV , 2021
2021
Cited alongside, same era.
Y. Chai, P. Sun, J. Ngiam, W. Wang, B. Caine, V. Vasudevan, X. Zhang, and D. Anguelov, “To the point: Efficient 3d object detection in the range image with graph convolution kernels,” in CVPR , 2021
2021
Cited alongside, same era.
J. Mao, Y. Xue, M. Niu, H. Bai, J. Feng, X. Liang, H. Xu, and C. Xu, “Voxel transformer for 3d object detection,” in CVPR , 2021
2021
Cited alongside, same era.
L. Fan, Z. Pang, T. Zhang, Y.-X. Wang, H. Zhao, F. Wang, N. Wang, and Z. Zhang, “Embracing single stride 3d object detector with sparse transformer,” arXiv preprint , 2021
2021
Cited alongside, same era.
Y. Wang and J. M. Solomon, “Object dgcnn: 3d object detection using dynamic graphs,” in NeurIPS , 2021
2021
Cited alongside, same era.
A. Piergiovanni, V. Casser, M. S. Ryoo, and A. Angelova, “4d-net for learned multi-modal alignment,” in CVPR , 2021
2021
Cited alongside, same era.
S. Casas, A. Sadat, and R. Urtasun, “Mp3: A unified model to map, perceive, predict and plan,” in CVPR , 2021
2021
Cited alongside, same era.
2022
Later among the works it cites.
T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Ré, “FlashAttention: Fast and memory-efficient exact attention with IO-awareness,” in NeurIPS , 2022
2022
Later among the works it cites.
L. Fan, F. Wang, N. Wang, and Z. Zhang, “Fully Sparse 3D Object Detection,” in NeurIPS , 2022
2022
Later among the works it cites.
T. Khurana, P. Hu, A. Dave, J. Ziglar, D. Held, and D. Ramanan, “Differentiable raycasting for self-supervised occupancy forecasting,” in ECCV , 2022
2022
Later among the works it cites.
S. Wang, Y. Liu, T. Wang, Y. Li, and X. Zhang, “Exploring object-centric temporal modeling for efficient multi-view 3d object detection,” in ICCV , 2023
2023
Later among the works it cites.
H. Liu, Y. Teng, T. Lu, H. Wang, and L. Wang, “Sparsebev: High-performance sparse 3d object detection from multi-camera videos,” in ICCV , 2023
2023
Later among the works it cites.
J. Li, C. Luo, and X. Yang, “Pillarnext: Rethinking network designs for 3d object detection in lidar point clouds,” in CVPR , 2023
2023
Later among the works it cites.
J. Huang, Y. Ye, Z. Liang, Y. Shan, and D. Du, “Detecting as labeling: Rethinking lidar-camera fusion in 3d object detection,” arXiv preprint , 2023
2023
Later among the works it cites.
H. Wang, H. Tang, S. Shi, A. Li, Z. Li, B. Schiele, and L. Wang, “Unitr: A unified and efficient multi-modal transformer for bird’s-eye-view representation,” in ICCV , 2023
2023
Later among the works it cites.
Y. Xie, C. Xu, M.-J. Rakotosaona, P. Rim, F. Tombari, K. Keutzer, M. Tomizuka, and W. Zhan, “Sparsefusion: Fusing multi-modal sparse representations for multi-sensor 3d object detection,” in ICCV , 2023
2023
Later among the works it cites.
Y. Hu, J. Yang, L. Chen, K. Li, C. Sima, X. Zhu, S. Chai, S. Du, T. Lin, W. Wang et al. , “Planning-oriented autonomous driving,” in CVPR , 2023
2023
Later among the works it cites.
B. Jiang, S. Chen, Q. Xu, B. Liao, J. Chen, H. Zhou, Q. Zhang, W. Liu, C. Huang, and X. Wang, “Vad: Vectorized scene representation for efficient autonomous driving,” in ICCV , 2023
2023
Later among the works it cites.
H. Liu, T. Lu, Y. Xu, J. Liu, and L. Wang, “Learning optical flow and scene flow with bidirectional camera-lidar fusion,” IEEE TPAMI , 2023
2023
Later among the works it cites.
T. Dao, “FlashAttention-2: Faster attention with better parallelism and work partitioning,” arXiv preprint , 2023
2023
Later among the works it cites.
J. Yan, Y. Liu, J. Sun, F. Jia, S. Li, T. Wang, and X. Zhang, “Cross modal transformer: Towards fast and robust 3d object detection,” in ICCV , 2023
2023
Later among the works it cites.
H. Wang, C. Shi, S. Shi, M. Lei, S. Wang, D. He, B. Schiele, and L. Wang, “Dsvt: Dynamic sparse voxel transformer with rotated sets,” in CVPR , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2024
Closest in time.
Y. Jiao, Z. Jie, S. Chen, J. Chen, L. Ma, and Y.-G. Jiang, “Msmdfusion: Fusing lidar and camera at multiple scales with multi-depth seeds for 3d object detection,” in CVPR , 2024
2024
Closest in time.
Z. Song, F. Jia, H. Pan, Y. Luo, C. Jia, G. Zhang, L. Liu, Y. Ji, L. Yang, and L. Wang, “Contrastalign: Toward robust bev feature alignment via contrastive learning for multi-modal 3d object detection,” arXiv preprint , 2024
2024
Closest in time.
Y. Li, L. Fan, Y. Liu, Z. Huang, Y. Chen, N. Wang, and Z. Zhang, “Fully sparse fusion for 3d object detection,” in IEEE TPAMI , 2024
2024
Closest in time.