Fetching the paper…
Reading the bibliography…
In recent years, vision-centric Bird's Eye View (BEV) perception has garnered significant interest from both industry and academia due to its inherent advantages, such as providing an intuitive representation of the world and being conducive to data fusion.
H. A. Mallot, H. H. Bülthoff, and J. e. Little, “Inverse perspective mapping simplifies optical flow computation and obstacle detection,” Biological cybernetics , vol. 64, no. 3, pp. 177–185, 1991
1991
Earlier work this paper cites.
R. Hartley and A. Zisserman, Multiple view geometry in computer vision . Cambridge university press, 2003
2003
Earlier work this paper cites.
T. Kim and T. Adali, “Approximation by fully complex multilayer perceptrons,” Neural Computation , vol. 15, pp. 1641–1666, 2003
2003
Earlier work this paper cites.
A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous driving? the kitti vision benchmark suite,” in CVPR , 2012
2012
Earlier work this paper cites.
S. Sengupta, P. Sturgess, L. Ladickỳ, and P. H. Torr, “Automatic dense visual semantic mapping from street-level imagery,” in IROS . IEEE, 2012, pp. 857–862
2012
Earlier work this paper cites.
A. Palazzi, G. Borghi, D. Abati, S. Calderara, and R. Cucchiara, “Learning to map vehicles into bird’s eye view,” in ICIAP , S. Battiato and G. G. etc., Eds., vol. 10484, 2017, pp. 233–243
2017
Earlier work this paper cites.
C. R. Qi, L. Yi, H. Su, and L. J. Guibas, “Pointnet++: Deep hierarchical feature learning on point sets in a metric space,” NeurIPS , vol. 30, 2017
2017
Earlier work this paper cites.
T.-Y. Lin, P. Dollár, R. Girshick, and K. e. He, “Feature pyramid networks for object detection,” in CVPR , 2017, pp. 2117–2125
2017
Earlier work this paper cites.
J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei, “Deformable convolutional networks,” ICCV , pp. 764–773, 2017
2017
Earlier work this paper cites.
A. Vaswani, N. M. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” ArXiv , 2017
2017
Earlier work this paper cites.
X. Chen, H. Ma, J. Wan, B. Li, and T. Xia, “Multi-view 3d object detection network for autonomous driving,” 2017, pp. 6526–6534
2017
Earlier work this paper cites.
A. Dosovitskiy, G. Ros, F. Codevilla, and A. M. L. etc., “Carla: An open urban driving simulator,” ArXiv , 2017
2017
Earlier work this paper cites.
J. Redmon and A. Farhadi, “Yolov3: an incremental improvement. arxiv [preprint] arxiv,” ArXiv , vol. 2, 2018
2018
Earlier work this paper cites.
A. Creswell, T. White, V. Dumoulin, K. Arulkumaran, and B. e. Sengupta, “Generative adversarial networks: An overview,” IEEE signal processing magazine , vol. 35, no. 1, pp. 53–65, 2018
2018
Earlier work this paper cites.
X. Zhu, Z. Yin, and J. e. Shi, “Generative adversarial frontal view to bird view synthesis,” in 3DV . IEEE, 2018, pp. 454–463
2018
Earlier work this paper cites.
T. Roddick, A. Kendall, and R. Cipolla, “Orthographic feature transform for monocular 3d object detection,” ArXiv , 2018
2018
Earlier work this paper cites.
D. Xu, D. Anguelov, and A. Jain, “Pointfusion: Deep sensor fusion for 3d bounding box estimation,” in CVPR , 2018, pp. 244–253
2018
Earlier work this paper cites.
M. Liang, B. Yang, and S. e. Wang, “Deep continuous fusion for multi-sensor 3d object detection,” in ECCV , 2018, pp. 641–656
2018
Earlier work this paper cites.
J. Ku, M. Mozifian, J. Lee, A. Harakeh, and S. L. Waslander, “Joint 3d proposal generation and object detection from view aggregation,” in IROS . IEEE, 2018, pp. 1–8
2018
Earlier work this paper cites.
C. R. Qi, W. Liu, C. Wu, and H. e. Su, “Frustum pointnets for 3d object detection from rgb-d data,” in CVPR , 2018, pp. 918–927
2018
Earlier work this paper cites.
Y. Yan, Y. Mao, and B. Li, “Second: Sparsely embedded convolutional detection,” Sensors , vol. 18, no. 10, 2018
2018
Earlier work this paper cites.
M.-F. Chang, J. Lambert, P. Sangkloy, and J. e. Singh, “Argoverse: 3d tracking and forecasting with rich maps,” in CVPR , 2019
2019
Earlier work this paper cites.
R. Kesten, M. Usman, J. Houston, T. Pandya, K. Nadhamuni, A. Ferreira, M. Yuan, B. Low, A. Jain, P. Ondruska, S. Omari, S. Shah, A. Kulkarni, and A. e. Kazakova, “Level 5 perception dataset 2020,” https://level-5.global/level5/data/
2019
Earlier work this paper cites.
A. Patil, S. Malla, H. Gang, and Y.-T. Chen, “The h3d dataset for full-surround 3d multi-object detection and tracking in crowded urban scenes,” in ICRA , 2019
2019
Earlier work this paper cites.
S. Ammar Abbas and A. Zisserman, “A geometric approach to obtain a bird’s eye view from an image,” in ICCVW , 2019, pp. 0–0
2019
Earlier work this paper cites.
Y. Kim and D. Kum, “Deep learning based vehicle position and orientation estimation via inverse perspective mapping image,” in IEEE Intelligent Vehicles Symposium . IEEE, 2019, pp. 317–323
2019
Earlier work this paper cites.
N. Garnett, R. Cohen, and T. e. Pe’er, “3d-lanenet: end-to-end 3d multiple lane detection,” in ICCV , 2019, pp. 2921–2930
2019
Earlier work this paper cites.
S. Srivastava, F. Jurie, and G. Sharma, “Learning 2d to 3d lifting for object detection in 3d for autonomous vehicles,” in IROS . IEEE, 2019, pp. 4504–4511
2019
Earlier work this paper cites.
T. Bruls, H. Porav, L. Kunze, and P. Newman, “The right (angled) perspective: Improving the understanding of road scenes using boosted inverse perspective mapping,” in IEEE Intelligent Vehicles Symposium . IEEE, 2019, pp. 302–309
2019
Earlier work this paper cites.
Y. Wang, W.-L. Chao, D. Garg, B. Hariharan, and M. e. Campbell, “Pseudo-lidar from visual depth estimation: Bridging the gap in 3d object detection for autonomous driving,” in CVPR , 2019
2019
Earlier work this paper cites.
X. Ma, Z. Wang, H. Li, P. Zhang, and W. e. Ouyang, “Accurate monocular 3d object detection via color-embedded 3d reconstruction for autonomous driving,” in ICCV , 2019, pp. 6851–6860
2019
Earlier work this paper cites.
C. Lu, M. J. G. van de Molengraft, and G. Dubbelman, “Monocular semantic occupancy grid mapping with convolutional variational encoder–decoder networks,” IEEE Robotics and Automation Letters , vol. 4, no. 2, pp. 445–452, 2019
2019
Earlier work this paper cites.
J. Lin, C. Gan, and S. Han, “Tsm: Temporal shift module for efficient video understanding,” in ICCV , 2019, pp. 7083–7093
2019
Earlier work this paper cites.
A. H. Lang, S. Vora, H. Caesar, and L. e. Zhou, “Pointpillars: Fast encoders for object detection from point clouds,” in CVPR , 2019
2019
Earlier work this paper cites.
B. Zhu and etc., “Class-balanced grouping and sampling for point cloud 3d object detection,” CoRR , 2019
2019
Earlier work this paper cites.
H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, and Y. e. Pan, “nuscenes: A multimodal dataset for autonomous driving,” in CVPR , 2020
2020
Earlier work this paper cites.
P. Sun, H. Kretzschmar, X. Dotiwalla, A. Chouard, V. Patnaik, P. Tsui, J. Guo, Y. Zhou, and Y. e. Chai, “Scalability in perception for autonomous driving: Waymo open dataset,” in ICCV , 2020
2020
Earlier work this paper cites.
N. Gählert, N. Jourdan, and M. e. Cordts, “Cityscapes 3d: Dataset and benchmark for 9 dof vehicle detection,” in CVPRW , 2020
2020
Earlier work this paper cites.
L. Reiher, B. Lampe, and L. Eckstein, “A sim2real deep learning approach for the transformation of images from multiple vehicle-mounted cameras to a semantically segmented image in bird’s eye view,” in ITSC . IEEE, 2020, pp. 1–7
2020
Earlier work this paper cites.
Y. Hou, L. Zheng, and S. Gould, “Multiview detection with feature perspective transformation,” in ECCV , ser. Lecture Notes in Computer Science, A. Vedaldi, H. Bischof, T. Brox, and J. Frahm, Eds., vol. 12352, 2020, pp. 1–18
2020
Earlier work this paper cites.
K. Mani, S. Daga, S. Garg, S. S. Narasimhan, M. Krishna, and K. M. Jatavallabhula, “Monolayout: Amodal scene layout from a single image,” in WACV , 2020, pp. 1689–1697
2020
Earlier work this paper cites.
Y. You, Y. Wang, W.-L. Chao, D. Garg, G. Pleiss, and B. e. Hariharan, “Pseudo-lidar++: Accurate depth for 3d object detection in autonomous driving,” in ICLR , 2020
2020
Earlier work this paper cites.
J. Philion and S. Fidler, “Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d,” in ECCV . Springer, 2020, pp. 194–210
2020
Earlier work this paper cites.
X. Ma, S. Liu, Z. Xia, and H. e. Zhang, “Rethinking pseudo-lidar representation,” in ECCV . Springer, 2020, pp. 311–327
2020
Earlier work this paper cites.
A. Simonelli, S. R. Bulò, L. Porzi, P. Kontschieder, and E. Ricci, “Are we missing confidence in pseudo-lidar methods for monocular 3d object detection?” ArXiv , 2020
2020
Earlier work this paper cites.
R. Qian, D. Garg, Y. Wang, Y. You, S. Belongie, B. Hariharan, M. Campbell, and W. etc., “End-to-end pseudo-lidar for image-based 3d object detection,” in CVPR , 2020, pp. 5881–5890
2020
Earlier work this paper cites.
Y. Chen, S. Liu, X. Shen, and J. Jia, “Dsgn: Deep stereo geometry network for 3d object detection,” in CVPR , 2020, pp. 12 536–12 545
2020
Earlier work this paper cites.
B. Pan, J. Sun, H. Y. T. Leung, A. Andonian, and B. Zhou, “Cross-view semantic segmentation for sensing surroundings,” IEEE Robotics and Automation Letters , vol. 5, no. 3, pp. 4867–4873, 2020
2020
Earlier work this paper cites.
N. Hendy, C. Sloan, F. Tian, P. Duan, N. Charchut, Y. Xie, C. Wang, and J. Philbin, “Fishing net: Future inference of semantic heatmaps in grids,” ArXiv , 2020
2020
Earlier work this paper cites.
T. Roddick and R. Cipolla, “Predicting semantic map representations from images using pyramid occupancy networks,” in CVPR , 2020, pp. 11 138–11 147
2020
Cited alongside, same era.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in ECCV . Springer, 2020, pp. 213–229
2020
Cited alongside, same era.
X. Zhu, W. Su, L. Lu, B. Li, and etc., “Deformable detr: Deformable transformers for end-to-end object detection,” ArXiv , 2020
2020
Cited alongside, same era.
N. Kitaev, L. Kaiser, and A. Levskaya, “Reformer: The efficient transformer,” ArXiv , 2020
2020
Cited alongside, same era.
Y. Tay, D. Bahri, L. Yang, D. Metzler, and D.-C. Juan, “Sparse sinkhorn attention,” in ICML , 2020
2020
Cited alongside, same era.
Y. B. Can and etc., “Topology preserving local road network estimation from single onboard camera image,” in CVPR , 2022, pp. 17 263–17 272
2022
Closest in time.
Y. Liu, J. Yan, F. Jia, and S. e. Li, “Petrv2: A unified framework for 3d perception from multi-camera images,” ArXiv , 2022
2022
Closest in time.
Z. Chen and Z. e. Li, “Graph-detr3d: Rethinking overlapping regions for multi-view 3d object detection,” ArXiv , 2022
2022
Closest in time.
W. K. Roh, G. Chang, S. Moon, G. Nam, C. Kim, Y. Kim, S. Kim, and J. Kim, “Ora3d: Overlap region aware multi-view 3d object detection,” ArXiv , 2022
2022
Closest in time.
S. Chen, , X. Wang, T. Cheng, and Q. e. Zhang, “Polar parametrization for vision-based surround-view 3d detection,” ArXiv , 2022
2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Wang, B. Z. Li, and M. K. etc., “Linformer: Self-attention with linear complexity,” ArXiv , 2020
2020
Cited alongside, same era.
S. Vora, A. H. Lang, B. Helou, and O. Beijbom, “Pointpainting: Sequential fusion for 3d object detection,” 2020, pp. 4603–4611
2020
Cited alongside, same era.
M. Zhu, S. Zhang, and Y. e. Zhong, “Monocular 3d vehicle detection using uncalibrated traffic cameras through homography,” in IROS , 2021, pp. 3814–3821
2021
Cited alongside, same era.
A. Loukkal and Y. e. Grandvalet, “Driving among flatmobiles: Bird-eye-view occupancy grids from a monocular camera for holistic trajectory planning,” in WACV , 2021, pp. 51–60
2021
Cited alongside, same era.
L. Song, J. Wu, M. Yang, Q. Zhang, Y. Li, and J. Yuan, “Stacked homography transformations for multi-view pedestrian detection,” in ICCV , 2021, pp. 6029–6037
2021
Cited alongside, same era.
X. Zhu, H. Zhou, T. Wang, F. Hong, W. Li, Y. Ma, and etc., “Cylindrical and asymmetrical 3d convolution networks for lidar-based perception,” TPAMI , vol. 44, no. 10, pp. 6807–6822, 2021
2021
Cited alongside, same era.
T. Yin, X. Zhou, and P. Krähenbühl, “Center-based 3d object detection and tracking,” CVPR , 2021
2021
Cited alongside, same era.
Y. Jiang, L. Zhang, Z. Miao, X. Zhu, J. Gao, W. Hu, and Y.-G. Jiang, “Polarformer: Multi-camera 3d object detection with polar transformer,” ArXiv , 2022
2022
Closest in time.
Y. Shi, J. Shen, Y. Sun, Y. Wang, J. Li, and S. S. etc., “Srcn3d: Sparse r-cnn 3d surround-view camera object detection and tracking for autonomous driving,” ArXiv , 2022
2022
Closest in time.
B. Zhou and P. Krähenbühl, “Cross-view transformers for real-time map-view semantic segmentation,” CoRR , 2022
2022
Closest in time.
L. Peng, Z. Chen, and Z. e. Fu, “Bevsegformer: Bird’s eye view semantic segmentation from arbitrary camera rigs,” ArXiv , 2022
2022
Closest in time.
L. Chen, C. Sima, Y. Li, Z. Zheng, J. Xu, X. Geng, H. Li, and C. e. He, “Persformer: 3d lane detection via perspective transformer and the openlane benchmark,” in ECCV , 2022
2022
Closest in time.
Z. Li, W. Wang, H. Li, E. Xie, C. Sima, T. Lu, Y. Qiao, and J. Dai, “Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers,” ArXiv , 2022
2022
Closest in time.
C. Yang, Y. Chen, H. Tian, C. Tao, X. Zhu, Z. Zhang, G. Huang, H. Li, Y. Qiao, L. Lu et al. , “Bevformer v2: Adapting modern image backbones to bird’s-eye-view recognition via perspective supervision,” Arxiv , 2022
2022
Closest in time.
J. Lu, Z. Zhou, X. Zhu, H. Xu, and L. Zhang, “Learning ego 3d representation as ray tracing,” ArXiv , 2022
2022
Closest in time.
S. Chen, T. Cheng, X. Wang, W. Meng, Q. Zhang, and W. Liu, “Efficient and robust 2d-to-bev representation learning via geometry-guided kernel transformer,” ArXiv , 2022
2022
Closest in time.
R. Xu, Z. Tu, H. Xiang, W. Shao, B. Zhou, and J. Ma, “Cobevt: Cooperative bird’s eye view semantic segmentation with sparse transformers,” ArXiv , 2022
2022
Closest in time.
A. Saha, O. Mendez, C. Russell, and R. Bowden, “Translating images into maps,” in ICRA . IEEE, 2022
2022
Closest in time.
S. Gong, X. Ye, X. Tan, J. Wang, E. Ding, Y. Zhou, and X. Bai, “Gitnet: Geometric prior-based transformation for birds-eye-view segmentation,” ArXiv , 2022
2022
Closest in time.
F. Bartoccioni, E. Zablocki, A. Bursuc, P. P’erez, M. Cord, and A. Karteek, “Lara: Latents and rays for multi-camera bird’s-eye-view semantic segmentation,” 2022
2022
Closest in time.
K.-C. Huang, T.-H. Wu, H.-T. Su, and W. H. Hsu, “Monodtr: Monocular 3d object detection with depth-aware transformer,” ArXiv , 2022
2022
Closest in time.
R. Zhang, H. Qiu, T. Wang, X. Xu, Z. Guo, Y. J. Qiao, P. Gao, and H. Li, “Monodetr: Depth-aware transformer for monocular 3d object detection,” ArXiv , 2022
2022
Closest in time.
A. K. Akan and F. Güney, “Stretchbev: Stretching future instance prediction spatially and temporally,” ArXiv , 2022
2022
Closest in time.
Z. Li, W. Wang, E. Xie, Z. Yu, A. Anandkumar, J. M. Alvarez, P. Luo, and T. Lu, “Panoptic segformer: Delving deeper into panoptic segmentation with transformers,” in CVPR , 2022, pp. 1280–1289
2022
Closest in time.
Y. Li, Y. Chen, and X. Q. etc., “Unifying voxel-based representation with transformer for 3d object detection,” CoRR , 2022
2022
Closest in time.
Z. Liu, H. Tang, A. Amini, X. Yang, H. Mao, D. Rus, and S. Han, “Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation,” CoRR , 2022
2022
Closest in time.
X. Chen, T. Zhang, Y. Wang, Y. Wang, and H. Zhao, “FUTR3D: A unified sensor fusion framework for 3d detection,” CoRR , 2022
2022
Closest in time.
Z. Chen, Z. Li, S. Zhang, L. Fang, Q. Jiang, F. Zhao, and B. Z. etc., “Autoalign: Pixel-instance feature aggregation for multi-modal 3d object detection,” in IJCAI , L. D. Raedt, Ed., 2022, pp. 827–833
2022
Closest in time.
Z. Chen, Z. Li, S. Zhang, L. Fang, Q. Jiang, and F. Zhao, “Autoalignv2: Deformable feature aggregation for dynamic multi-modal 3d object detection,” CoRR , 2022
2022
Closest in time.
A. W. Harley, Z. Fang, J. Li, R. Ambrus, and K. Fragkiadaki, “A simple baseline for BEV perception without lidar,” CoRR , 2022
2022
Closest in time.
X. Bai, Z. Hu, X. Zhu, Q. Huang, Y. Chen, H. Fu, and C. Tai, “Transfusion: Robust lidar-camera fusion for 3d object detection with transformers,” 2022
2022
Closest in time.
Z. Qin, J. Chen, C. Chen, X. Chen, and X. Li, “Uniformer: Unified multi-view fusion transformer for spatial-temporal representation in bird’s-eye-view,” ArXiv , 2022
2022
Closest in time.
“Tesla AI Day 2022,” 10 2022. [Online]. Available: https://www.youtube.com/watch?v=ODSJsviD_SU&t=2386s
2022
Closest in time.
A. Cao and R. de Charette, “Monoscene: Monocular 3d semantic scene completion,” in CVPR . IEEE, 2022, pp. 3981–3991
2022
Closest in time.
J. Huang and G. Huang, “Bevpoolv2: A cutting-edge implementation of bevdet toward deployment,” ArXiv , 2022
2022
Closest in time.
T. Wang and etc., “Probabilistic and geometric depth: Detecting objects in perspective,” in CoRL . PMLR, 2022, pp. 1475–1485
2022
Closest in time.
Y.-N. e. Chen, “Pseudo-stereo for monocular 3d object detection in autonomous driving,” in CVPR , June 2022, pp. 887–897
2022
Closest in time.
Y. Li, Z. Ge, G. Yu, J. Yang, and Z. e. Wang, “Bevdepth: Acquisition of reliable depth for multi-view 3d object detection,” ArXiv , 2022
2022
Closest in time.
C. Han, J. Sun, Z. Ge, J. Yang, R. Dong, H. Zhou, W. Mao, Y. Peng, and X. Zhang, “Exploring recurrent long-term temporal fusion for multi-view 3d perception,” Arxiv , 2023
2023
Closest in time.
Q. Lian, T. Wang, D. Lin, and J. Pang, “Dort: Modeling dynamic objects in recurrent for multi-camera 3d object detection and tracking,” Arxiv , 2023
2023
Closest in time.
J. He, Y. Chen, N. Wang, and Z. Zhang, “3d video object detection with learnable object-centric global optimization,” ArXiv , 2023
2023
Closest in time.
Z. Wang, Z. Huang, J. Fu, N. Wang, and S. Liu, “Object as query: Equipping any 2d object detector with 3d detection ability,” ArXiv , 2023
2023
Closest in time.
R. Miao, W. Liu, M. Chen, Z. Gong, W. Xu, C. Hu, and S. Zhou, “Occdepth: A depth-aware method for 3d semantic scene completion,” CoRR , 2023
2023
Closest in time.
Y. Huang, W. Zheng, Y. Zhang, J. Zhou, and J. Lu, “Tri-perspective view for vision-based 3d semantic occupancy prediction,” CoRR , 2023
2023
Closest in time.
Y. Li, Z. Yu, C. B. Choy, C. Xiao, J. M. Alvarez, S. Fidler, C. Feng, and A. Anandkumar, “Voxformer: Sparse voxel transformer for camera-based 3d semantic scene completion,” CoRR , 2023
2023
Closest in time.
Y. Zhang, Z. Zhu, and D. Du, “Occformer: Dual-path transformer for vision-based 3d semantic occupancy prediction,” 2023
2023
Closest in time.
Y. Wei, L. Zhao, W. Zheng, Z. Zhu, J. Zhou, and J. Lu, “Surroundocc: Multi-camera 3d occupancy prediction for autonomous driving,” CoRR , 2023
2023
Closest in time.
X. Wang, Z. Zhu, W. Xu, Y. Zhang, and etc., “Openoccupancy: A large scale benchmark for surrounding semantic occupancy perception,” CoRR , 2023
2023
Closest in time.
Z. Xia, Y. Liu, X. Li, X. Zhu, and Y. e. Ma, “Scpnet: Semantic scene completion on point cloud,” ArXiv , 2023
2023
Closest in time.
J. Liu, T. Wang, B. Liu, Q. Zhang, Y. Liu, and H. Li, “Towards better 3d knowledge transfer via masked image modeling for multi-view 3d understanding,” Arxiv , 2023
2023
Closest in time.
Z. Zong, D. Jiang, G. Song, Z. Xue, J. Su, H. Li, and Y. Liu, “Temporal enhanced training of multi-view 3d object detector via historical object prediction,” Arxiv , 2023
2023
Closest in time.