Fetching the paper…
Reading the bibliography…
Long-term temporal fusion is a crucial but often overlooked technique in camera-based Bird's-Eye-View (BEV) 3D perception.
T. Kanade and M. Okutomi, “A stereo matching algorithm with an adaptive window: Theory and experiment,” IEEE Trans. Pattern Anal. Mach. Intell. (TPAMI)
1994
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Comput
1997
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR)
2009
Earlier work this paper cites.
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei, “Large-scale video classification with convolutional neural networks,” in IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR)
2014
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” in Adv. Neural Inform. Process. Syst. (NIPS)
2014
Earlier work this paper cites.
K. Cho, B. van Merrienboer, D. Bahdanau, and Y. Bengio, “On the properties of neural machine translation: Encoder-decoder approaches,” in Empir. Method. Nat. Lang. Process. Worksh. (EMNLP Workshop)
2014
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in Adv. Neural Inform. Process. Syst. (NIPS)
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
M. Jaderberg, K. Simonyan, A. Zisserman, and K. Kavukcuoglu, “Spatial transformer networks,” in Adv. Neural Inform. Process. Syst. (NIPS)
2015
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Int. Conf. Medical Image Comput. Comput. Assist. Interv. (MICCAI)
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR)
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Adv. Neural Inform. Process. Syst. (NIPS)
2017
Earlier work this paper cites.
A. Mousavian, D. Anguelov, J. Flynn, and J. Kosecka, “3d bounding box estimation using deep learning and geometry,” in IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR)
2017
Earlier work this paper cites.
X. Chen, H. Ma, J. Wan, B. Li, and T. Xia, “Multi-view 3d object detection network for autonomous driving,” in IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR)
2017
Earlier work this paper cites.
W. Luo, B. Yang, and R. Urtasun, “Fast and furious: Real time end-to-end 3d detection, tracking and motion forecasting with a single convolutional net,” in IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR)
2018
Earlier work this paper cites.
Y. Yan, Y. Mao, and B. Li, “SECOND: sparsely embedded convolutional detection,” Sensors
2018
Earlier work this paper cites.
Z. Tian, C. Shen, H. Chen, and T. He, “FCOS: fully convolutional one-stage object detection,” in Int. Conf. Comput. Vis. (ICCV)
2019
Earlier work this paper cites.
A. H. Lang, S. Vora, H. Caesar, L. Zhou, J. Yang, and O. Beijbom, “Pointpillars: Fast encoders for object detection from point clouds,” in IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR)
2019
Earlier work this paper cites.
H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom, “nuscenes: A multimodal dataset for autonomous driving,” in IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR)
2020
Earlier work this paper cites.
J. Philion and S. Fidler, “Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d,” in Eur. Conf. Comput. Vis. (ECCV)
2020
Earlier work this paper cites.
X. Zhou, V. Koltun, and P. Krähenbühl, “Tracking objects as points,” in Eur. Conf. Comput. Vis. (ECCV)
2020
Earlier work this paper cites.
M. Liang, B. Yang, W. Zeng, Y. Chen, R. Hu, S. Casas, and R. Urtasun, “Pnpnet: End-to-end perception and prediction with tracking in the loop,” in IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR)
2020
Cited alongside, same era.
H. Song, W. Ding, Y. Chen, S. Shen, M. Y. Wang, and Q. Chen, “Pip: Planning-informed trajectory prediction for autonomous driving,” in Eur. Conf. Comput. Vis. (ECCV)
2020
Cited alongside, same era.
T. Wang, X. Zhu, J. Pang, and D. Lin, “FCOS3D: fully convolutional one-stage monocular 3d object detection,” in Int. Conf. Comput. Vis. Worksh. (ICCV Workshop)
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2022
Later among the works it cites.
2022
Later among the works it cites.
P. Li and J. Jin, “Time3d: End-to-end joint monocular 3d object detection and tracking for autonomous driving,” in IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR)
2022
Later among the works it cites.
N. Marinello, M. Proesmans, and L. V. Gool, “Triplettrack: 3d object tracking using triplet embeddings and LSTM,” in IEEE Conf. Comput. Vis. Pattern Recog. Worksh. (CVPR Workshop)
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Laddha, S. Gautam, G. P. Meyer, C. Vallespi-Gonzalez, and C. K. Wellington, “Rv-fusenet: Range view based fusion of time-series lidar data for joint 3d object detection and motion forecasting,” in IEEE/RSJ Int. Conf. Intell. Robot. and Syst. (IROS)
2021
Cited alongside, same era.
T. Yin, X. Zhou, and P. Krähenbühl, “Center-based 3d object detection and tracking,” in IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR)
2021
Cited alongside, same era.
T. Wang, X. Zhu, J. Pang, and D. Lin, “FCOS3D: fully convolutional one-stage monocular 3d object detection,” in Int. Conf. Comput. Vis. Worksh. (ICCV Workshop)
2021
Cited alongside, same era.
Y. Wang, V. Guizilini, T. Zhang, Y. Wang, H. Zhao, and J. Solomon, “DETR3D: 3d object detection from multi-view images via 3d-to-2d queries,” in Annu. Conf. Robot. Learn. (CoRL)
2021
Cited alongside, same era.
A. Hu, Z. Murez, N. Mohan, S. Dudas, J. Hawke, V. Badrinarayanan, R. Cipolla, and A. Kendall, “FIERY: future instance prediction in bird’s-eye view from surround monocular cameras,” in Int. Conf. Comput. Vis. (ICCV)
2021
Cited alongside, same era.
2021
Cited alongside, same era.
J. Huang and G. Huang, “Bevdet4d: Exploit temporal cues in multi-camera 3d object detection,” CoRR
2022
Cited alongside, same era.
2022
Cited alongside, same era.
T. Zhang, X. Chen, Y. Wang, Y. Wang, and H. Zhao, “MUTR3D: A multi-camera tracking framework via 3d-to-2d queries,” in IEEE Conf. Comput. Vis. Pattern Recog. Worksh. (CVPR Workshop)
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
N. Peri, J. Luiten, M. Li, A. Osep, L. Leal-Taixé, and D. Ramanan, “Forecasting from lidar via future object detection,” in IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR)
2022
Later among the works it cites.
Y. Li, Z. Ge, G. Yu, J. Yang, Z. Wang, Y. Shi, J. Sun, and Z. Li, “Bevdepth: Acquisition of reliable depth for multi-view 3d object detection,” in AAAI Conf. Artif. Intell. (AAAI)
2023
Closest in time.
J. Park, C. Xu, S. Yang, K. Keutzer, K. Kitani, M. Tomizuka, and W. Zhan, “Time will tell: New outlooks and A baseline for temporal multi-view 3d object detection,” in Int. Conf. Learn. Represent. (ICLR)
2023
Closest in time.
C. Yang, Y. Chen, H. Tian, C. Tao, X. Zhu, Z. Zhang, G. Huang, H. Li, Y. Qiao, L. Lu, J. Zhou, and J. Dai, “Bevformer v2: Adapting modern image backbones to bird’s-eye-view recognition via perspective supervision,” in IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR)
2023
Closest in time.
Y. Li, H. Bao, Z. Ge, J. Yang, J. Sun, and Z. Li, “Bevstereo: Enhancing depth estimation in multi-view 3d object detection with dynamic temporal stereo,” in AAAI Conf. Artif. Intell. (AAAI)
2023
Closest in time.
Z. Zong, D. Jiang, G. Song, Z. Xue, J. Su, H. Li, and Y. Liu, “Temporal enhanced training of multi-view 3d object detector via historical object prediction,” in Int. Conf. Comput. Vis. (ICCV)
2023
Closest in time.
Z. Wang, C. Min, Z. Ge, Y. Li, Z. Li, H. Yang, and D. Huang, “STS: surround-view temporal stereo for multi-view 3d detection,” in AAAI Conf. Artif. Intell. (AAAI)
2023
Closest in time.
Y. Jiang, L. Zhang, Z. Miao, X. Zhu, J. Gao, W. Hu, and Y. Jiang, “Polarformer: Multi-camera 3d object detection with polar transformers,” in AAAI Conf. Artif. Intell. (AAAI)
2023
Closest in time.
W. Wang, J. Dai, Z. Chen, Z. Huang, Z. Li, X. Zhu, X. Hu, T. Lu, L. Lu, H. Li, X. Wang, and Y. Qiao, “Internimage: Exploring large-scale vision foundation models with deformable convolutions,” in IEEE/CVF Conf. Comput. Vis. Pattern Recog. (CVPR)
2023
Closest in time.
H. Hu, Y. Yang, T. Fischer, T. Darrell, F. Yu, and M. Sun, “Monocular quasi-dense 3d object tracking,” IEEE Trans. Pattern Anal. Mach. Intell. (TPAMI)
2023
Closest in time.