Fetching the paper…
Reading the bibliography…
Vision-based Bird's Eye View (BEV) representation is an emerging perception formulation for autonomous driving.
D. Scharstein and R. Szeliski, “A taxonomy and evaluation of dense two-frame stereo correspondence algorithms,” IJCV , vol. 47, no. 1, pp. 7–42, 2002
2002
Earlier work this paper cites.
H. Fu, M. Gong, C. Wang, K. Batmanghelich, and D. Tao, “Deep ordinal regression network for monocular depth estimation,” in CVPR , 2018, pp. 2002–2011
2011
Earlier work this paper cites.
D. Eigen, C. Puhrsch, and R. Fergus, “Depth map prediction from a single image using a multi-scale deep network,” NeurIPS , vol. 27, 2014
2014
Earlier work this paper cites.
J. Flynn, I. Neulander, J. Philbin, and N. Snavely, “Deepstereo: Learning to predict new views from the world’s imagery,” in CVPR , 2016, pp. 5515–5524
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” NeurIPS , vol. 30, 2017
2017
Earlier work this paper cites.
A. Kendall and Y. Gal, “What uncertainties do we need in bayesian deep learning for computer vision?” NeurIPS , vol. 30, 2017
2017
Earlier work this paper cites.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in ICCV , 2017, pp. 2980–2988
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
B. Xu and Z. Chen, “Multi-level fusion based 3d object detection from monocular images,” in CVPR , 2018, pp. 2345–2353
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
Y. Wang, W.-L. Chao, D. Garg, B. Hariharan, M. Campbell, and K. Q. Weinberger, “Pseudo-lidar from visual depth estimation: Bridging the gap in 3d object detection for autonomous driving,” in CVPR , 2019, pp. 8445–8453
2019
Earlier work this paper cites.
A. H. Lang, S. Vora, H. Caesar, L. Zhou, J. Yang, and O. Beijbom, “Pointpillars: Fast encoders for object detection from point clouds,” in CVPR , 2019, pp. 12 697–12 705
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Y. Lee, J.-w. Hwang, S. Lee, Y. Bae, and J. Park, “An energy and gpu-computation efficient backbone network for real-time object detection,” in CVPRW , 2019
2019
Earlier work this paper cites.
J. Philion and S. Fidler, “Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d,” in ECCV , 2020, pp. 194–210
2020
Earlier work this paper cites.
Y. You, Y. Wang, W.-L. Chao, D. Garg, G. Pleiss, B. Hariharan, M. Campbell, and K. Q. Weinberger, “Pseudo-lidar++: Accurate depth for 3d object detection in autonomous driving,” in ICLR , 2020
2020
Earlier work this paper cites.
X. Ma, S. Liu, Z. Xia, H. Zhang, X. Zeng, and W. Ouyang, “Rethinking pseudo-lidar representation,” in ECCV , 2020, pp. 311–327
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” NeurIPS , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in ECCV , 2020, pp. 213–229
2020
Cited alongside, same era.
M. Ding, Y. Huo, H. Yi, Z. Wang, J. Shi, Z. Lu, and P. Luo, “Learning depth-guided convolutions for monocular 3d object detection,” in CVPRW , 2020, pp. 1000–1001
2020
Cited alongside, same era.
Y. Cai, B. Li, Z. Jiao, H. Li, X. Zeng, and X. Wang, “Monocular 3d object detection with decoupled structured polygon estimation and height-guided depth estimation,” in AAAI , vol. 34, 2020, pp. 10 478–10 485
2020
Cited alongside, same era.
Y. Chen, L. Tai, K. Sun, and M. Li, “Monopair: Monocular 3d object detection using pairwise spatial relationships,” in CVPR , 2020, pp. 12 093–12 102
2020
Cited alongside, same era.
P. Lu, S. Xu, and H. Peng, “Graph-embedded lane detection,” IEEE TIP , vol. 30, pp. 2977–2988, 2021
2022
Later among the works it cites.
Z. Qin and X. Li, “Monoground: Detecting monocular 3d objects from the ground,” in CVPR , 2022, pp. 3793–3802
2022
Later among the works it cites.
T. Wang, X. Zhu, J. Pang, and D. Lin, “Probabilistic and geometric depth: Detecting objects in perspective,” in CoRL , 2022, pp. 1475–1485
2022
Later among the works it cites.
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
D. Park, R. Ambrus, V. Guizilini, J. Li, and A. Gaidon, “Is pseudo-lidar needed for monocular 3d object detection?” in ICCV , 2021, pp. 3142–3152
2021
Cited alongside, same era.
T. Yin, X. Zhou, and P. Krahenbuhl, “Center-based 3d object detection and tracking,” in CVPR , 2021, pp. 11 784–11 793
2021
Cited alongside, same era.
C. Reading, A. Harakeh, J. Chae, and S. L. Waslander, “Categorical depth distribution network for monocular 3d object detection,” in CVPR , 2021, pp. 8555–8564
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in ICCV , 2021, pp. 10 012–10 022
2021
Cited alongside, same era.
Y. Zhang, J. Lu, and J. Zhou, “Objects are different: Flexible monocular 3d object detection,” in CVPR , 2021, pp. 3289–3298
2021
Cited alongside, same era.
Y. Lu, X. Ma, L. Yang, T. Zhang, Y. Liu, Q. Chu, J. Yan, and W. Ouyang, “Geometry uncertainty projection network for monocular 3d object detection,” in ICCV , 2021, pp. 3111–3121
2021
Cited alongside, same era.
2022
Later among the works it cites.
Y. Li, Y. Chen, X. Qi, Z. Li, J. Sun, and J. Jia, “Unifying voxel-based representation with transformer for 3d object detection,” in NeurIPS , 2022
2022
Later among the works it cites.
J. Lu, Z. Zhou, X. Zhu, H. Xu, and L. Zhang, “Learning ego 3d representation as ray tracing,” in ECCV , 2022
2022
Later among the works it cites.
L. Xie, G. Xu, D. Cai, and X. He, “X-view: non-egocentric multi-view 3d object detector,” IEEE TIP , vol. 32, pp. 1488–1497, 2023
2023
Closest in time.
C. Yang, Y. Chen, H. Tian, C. Tao, X. Zhu, Z. Zhang, G. Huang, H. Li, Y. Qiao, L. Lu et al. , “Bevformer v2: Adapting modern image backbones to bird’s-eye-view recognition via perspective supervision,” in CVPR , 2023, pp. 17 830–17 839
2023
Closest in time.
F. Bartoccioni, É. Zablocki, A. Bursuc, P. Pérez, M. Cord, and K. Alahari, “Lara: Latents and rays for multi-camera bird’s-eye-view semantic segmentation,” in CoRL , 2023, pp. 1663–1672
2023
Closest in time.
Y. Li, Z. Ge, G. Yu, J. Yang, Z. Wang, Y. Shi, J. Sun, and Z. Li, “Bevdepth: Acquisition of reliable depth for multi-view 3d object detection,” in AAAI , vol. 37, no. 2, 2023, pp. 1477–1485
2023
Closest in time.
Y. Liu, J. Yan, F. Jia, S. Li, A. Gao, T. Wang, and X. Zhang, “Petrv2: A unified framework for 3d perception from multi-camera images,” in ICCV , 2023, pp. 3262–3272
2023
Closest in time.
Y. Jiang, L. Zhang, Z. Miao, X. Zhu, J. Gao, W. Hu, and Y.-G. Jiang, “Polarformer: Multi-camera 3d object detection with polar transformer,” in AAAI , vol. 37, no. 1, 2023, pp. 1042–1050
2023
Closest in time.
S. Fang, Z. Wang, Y. Zhong, J. Ge, and S. Chen, “Tbp-former: Learning temporal bird’s-eye-view pyramid for joint perception and prediction in vision-centric autonomous driving,” in CVPR , 2023, pp. 1368–1378
2023
Closest in time.
2023
Closest in time.
Y. Li, B. Huang, Z. Chen, Y. Cui, F. Liang, M. Shen, F. Liu, E. Xie, L. Sheng, W. Ouyang, and J. Shao, “Fast-bev: A fast and strong bird’s-eye view perception baseline,” IEEE TPAMI , pp. 1–14, 2024
2024
Closest in time.
W. Liu, Q. Li, W. Yang, J. Cai, Y. Yu, Y. Ma, S. He, and J. Pan, “Monocular bev perception of road scenes via front-to-top view projection,” IEEE TPAMI , pp. 1–17, 2024
2024
Closest in time.