Fetching the paper…
Reading the bibliography…
The field of autonomous driving increasingly demands high-quality annotated video training data.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Earlier work this paper cites.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” in
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom, “nuscenes: A multimodal dataset for autonomous driving,” in
2020
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in
2020
Earlier work this paper cites.
B. Sheng, Y. Fang, F. Xiao, and L. Sun, “An accurate device-free action recognition system using two-stream network,”
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
P. Dhariwal and A. Q. Nichol, “Diffusion models beat gans on image synthesis,” in
2021
Earlier work this paper cites.
Y. Wang, V. Guizilini, T. Zhang, Y. Wang, H. Zhao, and J. Solomon, “DETR3D: 3d object detection from multi-view images via 3d-to-2d queries,” in
2021
Earlier work this paper cites.
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” in
2021
Earlier work this paper cites.
G. Bertasius, H. Wang, and L. Torresani, “Is space-time attention all you need for video understanding?” in
2021
Earlier work this paper cites.
A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lucic, and C. Schmid, “Vivit: A video vision transformer,” in
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in
2021
Earlier work this paper cites.
Z. Li, W. Wang, H. Li, E. Xie, C. Sima, T. Lu, Y. Qiao, and J. Dai, “Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers,” in
2022
Earlier work this paper cites.
L. Chen, C. Sima, Y. Li, Z. Zheng, J. Xu, X. Geng, H. Li, C. He, J. Shi, Y. Qiao
2022
Earlier work this paper cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in
2022
Earlier work this paper cites.
A. Q. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen, “GLIDE: towards photorealistic image generation and editing with text-guided diffusion models,” in
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, S. K. S. Ghasemipour, R. G. Lopes, B. K. Ayan, T. Salimans, J. Ho, D. J. Fleet, and M. Norouzi, “Photorealistic text-to-image diffusion models with deep language understanding,” in
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Y. Liu, T. Wang, X. Zhang, and J. Sun, “Petr: Position embedding transformation for multi-view 3d object detection,” in
2022
Cited alongside, same era.
T. Zhang, X. Chen, Y. Wang, Y. Wang, and H. Zhao, “MUTR3D: A multi-camera tracking framework via 3d-to-2d queries,” in
2022
Cited alongside, same era.
Z. Lin, S. Geng, R. Zhang, P. Gao, G. de Melo, X. Wang, J. Dai, Y. Qiao, and H. Li, “Frozen CLIP models are efficient video learners,” in
2022
Cited alongside, same era.
B. Zhou and P. Krähenbühl, “Cross-view transformers for real-time map-view semantic segmentation,” in
2022
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Liu, J. Yan, F. Jia, S. Li, A. Gao, T. Wang, and X. Zhang, “Petrv2: A unified framework for 3d perception from multi-camera images,” in
2023
Cited alongside, same era.
Y. Jiang, L. Zhang, Z. Miao, X. Zhu, J. Gao, W. Hu, and Y.-G. Jiang, “Polarformer: Multi-camera 3d object detection with polar transformer,” in
2023
Cited alongside, same era.
Z. Pang, J. Li, P. Tokmakov, D. Chen, S. Zagoruyko, and Y. Wang, “Standing between past and future: Spatio-temporal modeling for multi-camera 3d multi-object tracking,” in
2023
Cited alongside, same era.
S. Huang, Z. Shen, Z. Huang, Z.-h. Ding, J. Dai, J. Han, N. Wang, and S. Liu, “Anchor3dlane: Learning to regress 3d anchors for monocular 3d lane detection,” in
2023
Cited alongside, same era.
2023
Cited alongside, same era.
A. Swerdlow, R. Xu, and B. Zhou, “Street-view image generation from a bird’s-eye view layout,”
2023
Cited alongside, same era.
G. Zheng, X. Zhou, X. Li, Z. Qi, Y. Shan, and X. Li, “Layoutdiffusion: Controllable diffusion model for layout-to-image generation,” in
2023
Cited alongside, same era.
2023
Later among the works it cites.
A. Blattmann, R. Rombach, H. Ling, T. Dockhorn, S. W. Kim, S. Fidler, and K. Kreis, “Align your latents: High-resolution video synthesis with latent diffusion models,” in
2023
Later among the works it cites.
U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni, D. Parikh, S. Gupta, and Y. Taigman, “Make-a-video: Text-to-video generation without text-video data,” in
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Li, Z. Ge, G. Yu, J. Yang, Z. Wang, Y. Shi, J. Sun, and Z. Li, “Bevdepth: Acquisition of reliable depth for multi-view 3d object detection,” in
2023
Later among the works it cites.
Y. Hu, J. Yang, L. Chen, K. Li, C. Sima, X. Zhu, S. Chai, S. Du, T. Lin, W. Wang, L. Lu, X. Jia, Q. Liu, J. Dai, Y. Qiao, and H. Li, “Planning-oriented autonomous driving,” in
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Park, C. Xu, S. Yang, K. Keutzer, K. M. Kitani, M. Tomizuka, and W. Zhan, “Time will tell: New outlooks and A baseline for temporal multi-view 3d object detection,” in
2023
Later among the works it cites.
Y. Zhao, C. Luo, C. Tang, D. Chen, N. C. Codella, L. Yuan, and Z.-J. Zha, “T2d: Spatiotemporal feature learning based on triple 2d decomposition,” 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
D. Chen, J. Li, V. Guizilini, R. Ambrus, and A. Gaidon, “Viewpoint equivariance for multi-view 3d object detection,” in
2023
Later among the works it cites.
B. Liao, S. Chen, X. Wang, T. Cheng, Q. Zhang, W. Liu, and C. Huang, “Maptr: Structured modeling and learning for online vectorized HD map construction,” in
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Que, L. Xiong, W. Wan, X. Xia, and Z. Liu, “Denoising diffusion probabilistic model for face sketch-to-photo synthesis,”
2024
Closest in time.
X. Jiang, S. Li, Y. Liu, S. Wang, F. Jia, T. Wang, L. Han, and X. Zhang, “Far3d: Expanding the horizon for surround-view 3d object detection,” in
2024
Closest in time.
2024
Closest in time.