Fetching the paper…
Reading the bibliography…
3D object detection plays a pivotal role in autonomous driving and robotics, demanding precise interpretation of Bird's Eye View (BEV) images.
The Hungarian method for the assignment problem
Harold W Kuhn. 1955 · 1955
Earlier work this paper cites.
Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Focal loss for dense object detection. In Proceedings of the IEEE international conference on computer vision . 2980–2988
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Pseudo-lidar from visual depth estimation: Bridging the gap in 3d object detection for autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 8445–8453
Yan Wang, Wei-Lun Chao, Divyansh Garg, Bharath Hariharan, Mark Campbell, and Kilian Q Weinberger. 2019 · 2019
Earlier work this paper cites.
nuscenes: A multimodal dataset for autonomous driving. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 11621–11631
Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. 2020 · 2020
Earlier work this paper cites.
End-to-end object detection with transformers. In Proceedings of the European conference on computer vision . 213–229
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020 · 2020
Earlier work this paper cites.
Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XIV 16 . Springer, 194–210
Jonah Philion and Sanja Fidler. 2020 · 2020
Earlier work this paper cites.
Scalability in perception for autonomous driving: Waymo open dataset. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 2446–2454
Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al · 2020
Earlier work this paper cites.
Deformable detr: Deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. 2020 · 2020
Earlier work this paper cites.
Bevdet: High-performance multi-camera 3d object detection in bird-eye-view
Junjie Huang, Guan Huang, Zheng Zhu, and Dalong Du. 2021 · 2021
Earlier work this paper cites.
Categorical depth distribution network for monocular 3d object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 8555–8564
Cody Reading, Ali Harakeh, Julia Chae, and Steven L Waslander. 2021 · 2021
Earlier work this paper cites.
Sparse r-cnn: End-to-end object detection with learnable proposals. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 14454–14463
Peize Sun, Rufeng Zhang, Yi Jiang, Tao Kong, Chenfeng Xu, Wei Zhan, Masayoshi Tomizuka, Lei Li, Zehuan Yuan, Changhu Wang, et al · 2021
Cited alongside, same era.
Fcos3d: Fully convolutional one-stage monocular 3d object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . 913–922
Tai Wang, Xinge Zhu, Jiangmiao Pang, and Dahua Lin. 2021 · 2021
Cited alongside, same era.
AdaMixer: A Fast-Converging Query-Based Object Detector. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5364–5373
Ziteng Gao, Limin Wang, Bing Han, and Sheng Guo. 2022 · 2022
Cited alongside, same era.
Bevdet4d: Exploit temporal cues in multi-camera 3d object detection
Junjie Huang and Guan Huang. 2022a · 2022
Cited alongside, same era.
Time will tell: New outlooks and a baseline for temporal multi-view 3d object detection
Jinhyung Park, Chenfeng Xu, Shijia Yang, Kurt Keutzer, Kris Kitani, Masayoshi Tomizuka, and Wei Zhan. 2022 · 2022
Later among the works it cites.
Mv-fcos3d++: Multi-view camera-only 4d object detection with pretrained monocular backbones
Tai Wang, Qing Lian, Chenming Zhu, Xinge Zhu, and Wenwei Zhang. 2022b · 2022
Later among the works it cites.
Mrrl: Modifying the reference via reinforcement learning for non-autoregressive joint multiple intent detection and slot filling. In Findings of the Association for Computational Linguistics: EMNLP 2023 . 10495–10505
Xuxin Cheng, Zhihong Zhu, Bowen Cao, Qichen Ye, and Yuexian Zou. 2023a · 2023
Closest in time.
Accelerating multiple intent detection and slot filling via targeted knowledge distillation. In The 2023 Conference on Empirical Methods in Natural Language Processing
Xuxin Cheng, Zhihong Zhu, Wanshi Xu, Yaowei Li, Hongxiang Li, and Yuexian Zou. 2023b · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Junjie Huang and Guan Huang. 2022b · 2022
Cited alongside, same era.
BEVStereo: Enhancing Depth Estimation in Multi-view 3D Object Detection with Dynamic Temporal Stereo
Yinhao Li, Han Bao, Zheng Ge, Jinrong Yang, Jianjian Sun, and Zeming Li. 2022a · 2022
Cited alongside, same era.
Unifying Voxel-based Representation with Transformer for 3D Object Detection
Y. Li, Y. Chen, X. Qi, Z. Li, J. Sun, and J. Jia. 2022b · 2022
Cited alongside, same era.
Bevdepth: Acquisition of reliable depth for multi-view 3d object detection
Yinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang, Zengran Wang, Yukang Shi, Jianjian Sun, and Zeming Li. 2022c · 2022
Cited alongside, same era.
BEVFormer ++ : Improving BEVFormer for 3D Camera-only Object Detection: 1st Place Solution for Waymo Open Dataset Challenge 2022
Zhiqi Li, Hanming Deng, Tianyu Li, Yangyi Huang, Chonghao Sima, Xiangwei Geng, Yulu Gao, Wenhai Wang, Yang Li, and Lewei Lu. 2023 · 2022
Cited alongside, same era.
Sparse4d: Multi-view 3d object detection with sparse spatial-temporal fusion
Xuewu Lin, Tianwei Lin, Zixiang Pei, Lichao Huang, and Zhizhong Su. 2022 · 2022
Cited alongside, same era.
PETRv2: A Unified Framework for 3D Perception from Multi-Camera Images
Yingfei Liu, Junjie Yan, Fan Jia, Shuailin Li, Qi Gao, Tiancai Wang, Xiangyu Zhang, and Jian Sun. 2022b · 2022
Cited alongside, same era.
BEVFormer: Learning Bird’s-Eye-View Representation from Multi-camera Images via Spatiotemporal Transformers. In European Conference on Computer Vision (ECCV) . Springer, 1–18
Zhiqi Li, Wenhai Wang, Hongyang Li, Enze Xie, Chonghao Sima, Tong Lu, Yu Qiao, and Jifeng Dai. 2022d
Cited in the paper.
Sparse4D v2: Recurrent Temporal Fusion with Sparse Model
Xuewu Lin, Tianwei Lin, Zixiang Pei, Lichao Huang, and Zhizhong Su. 2023 · 2023
Closest in time.
SparseBEV: High-Performance Sparse 3D Object Detection from Multi-Camera Videos
Haisong Liu, Yao Teng, Tao Lu, Haiguang Wang, and Limin Wang. 2023 · 2023
Closest in time.
Exploring Object-Centric Temporal Modeling for Efficient Multi-View 3D Object Detection
Shihao Wang, Yingfei Liu, Tiancai Wang, Ying Li, and Xiangyu Zhang. 2023b · 2023
Closest in time.
Cross Modal Transformer: Towards Fast and Robust 3D Object Detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 18268–18278
J. Yan, Y. Liu, J. Sun, F. Jia, S. Li, T. Wang, and X. Zhang. 2023 · 2023
Closest in time.
Ndc-scene: Boost monocular 3d semantic scene completion in normalized device coordinates space. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV) . IEEE Computer Society, 9421–9431
Jiawei Yao, Chuming Li, Keqiang Sun, Yingjie Cai, Hao Li, Wanli Ouyang, and Hongsheng Li. 2023 · 2023
Closest in time.
A Simple Baseline for Multi-camera 3D Object Detection. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 37. 3507–3515
Y. Zhang, W. Zheng, Z. Zhu, G. Huang, J. Lu, and J. Zhou. 2023 · 2023
Closest in time.
Building lane-level maps from aerial images. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 3890–3894
Jiawei Yao, Xiaochao Pan, Tong Wu, and Xiaofeng Zhang. 2024 · 2024
Closest in time.