Fetching the paper…
Reading the bibliography…
Accurate and high-fidelity driving scene reconstruction demands the effective utilization of comprehensive scene information as conditional inputs.
T. Whitted, “An improved illumination model for shaded display,” in ACM Siggraph 2005 Courses , 2005, pp. 4–es
2005
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, part III 18 . Springer, 2015, pp. 234–241
2015
Earlier work this paper cites.
I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning . MIT Press, 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
B. Dhingra, H. Liu, Z. Yang, W. W. Cohen, and R. Salakhutdinov, “Gated-attention readers for text comprehension,” in Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2017, pp. 1832–1846
2017
Earlier work this paper cites.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” in Advances in Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
M. Martínez-Díaz and F. Soriguera, “Autonomous vehicles: theoretical and practical challenges,” Transportation Research Procedia , vol. 33, pp. 275–282, 2018, xIII Conference on Transport Engineering, CIT2018. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S2352146518302606
2018
Earlier work this paper cites.
P. M. Bösch, F. Becker, H. Becker, and K. W. Axhausen, “Cost-based analysis of autonomous mobility services,” Transport Policy , vol. 64, pp. 76–91, 2018. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0967070X17300811
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
T. Unterthiner, B. Nessler, G. Heigold, S. Szedmak, and S. Hochreiter, “Towards accurate generative models of video: A new metric and challenges,” in Workshop on Challenges and Opportunities for AI in Financial Services at NeurIPS , 2018
2018
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom, “nuscenes: A multimodal dataset for autonomous driving,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 11 621–11 631
2020
Earlier work this paper cites.
P. Sun, H. Kretzschmar, X. Dotiwalla, A. Chouard, V. Patnaik, P. Tsui, J. Guo, Y. Zhou, Y. Chai, B. Caine, V. Vasudevan, W. Han, J. Ngiam, H. Zhao, A. Timofeev, S. Ettinger, M. Krivokon, A. Gao, A. Joshi, Y. Zhang, J. Shlens, Z. Chen, and D. Anguelov, “Scalability in perception for autonomous driving: Waymo open dataset,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2020
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” Communications of the ACM , vol. 65, no. 1, pp. 99–106, 2021
2021
Earlier work this paper cites.
P. Esser, R. Rombach, and B. Ommer, “Taming transformers for high-resolution image synthesis,” pp. 12 873–12 883, 2021
2021
Earlier work this paper cites.
S. W. Kim, J. Philion, A. Torralba, and S. Fidler, “Drivegan: Towards a controllable high-quality neural simulation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 5820–5829
2021
Cited alongside, same era.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 10 684–10 695
2022
Cited alongside, same era.
U. Singer, “Make-a-video: Text-to-video generation without text-video data,” 2022. [Online]. Available: https://makeavideo.studio/Make-A-Video.pdf
2022
Cited alongside, same era.
B. Zhou and P. Krähenbühl, “Cross-view transformers for real-time map-view semantic segmentation,” in CVPR , 2022
2022
Cited alongside, same era.
J. Li, D. Li, J. Gao et al. , “Blip-2: Bootstrapping language-image pretraining with frozen vision-language models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2023
2023
Later among the works it cites.
C. Cui, Y. Ma, X. Cao, W. Ye, Y. Zhou, K. Liang, J. Chen, J. Lu, Z. Yang, K.-D. Liao et al. , “A survey on multimodal large language models for autonomous driving,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2024, pp. 958–979
2024
Later among the works it cites.
L. Chen, P. Wu, K. Chitta, B. Jaeger, A. Geiger, and H. Li, “End-to-end autonomous driving: Challenges and frontiers,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2024
2024
Later among the works it cites.
L. Szabó and Z. Weltsch, “A comprehensive review of existing datasets for off-road autonomous vehicles,” in 2024 IEEE 22nd World Symposium on Applied Machine Intelligence and Informatics (SAMI) , 2024, pp. 000 403–000 410
2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Li, W. Wang, H. Li, E. Xie, C. Sima, T. Lu, Y. Qiao, and J. Dai, “Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers,” in European conference on computer vision . Springer, 2022, pp. 1–18
2022
Cited alongside, same era.
J. Ho, X. Chen, A. Srinivas, and et al., “Classifier-free diffusion guidance,” in NeurIPS 2022 , 2022
2022
Cited alongside, same era.
2023
Cited alongside, same era.
R. Gao, K. Chen, E. Xie, L. Hong, Z. Li, D.-Y. Yeung, and Q. Xu, “Magicdrive: Street view generation with diverse 3d geometry control,” in International Conference on Learning Representations , 2023
2023
Cited alongside, same era.
L. Zhang, A. Rao, and M. Agrawala, “Adding conditional control to text-to-image diffusion models,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 3836–3847
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Later among the works it cites.
Y. Wen, Y. Zhao, Y. Liu, F. Jia, Y. Wang, C. Luo, C. Zhang, T. Wang, X. Sun, and X. Zhang, “Panacea: Panoramic and controllable video generation for autonomous driving,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 6902–6912
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Wang, J. He, L. Fan, H. Li, Y. Chen, and Z. Zhang, “Driving into the future: Multiview visual forecasting and planning with world model for autonomous driving,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 14 749–14 759
2024
Later among the works it cites.
S. Zhao, D. Chen, Y.-C. Chen, J. Bao, S. Hao, L. Yuan, and K.-Y. K. Wong, “Uni-controlnet: All-in-one control to text-to-image diffusion models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
CODA Dataset, “w-coda 2024 track 2,” 2024, accessed: 2024-01-07. [Online]. Available: https://coda-dataset.github.io/w-coda2024/track2/
2024
Later among the works it cites.
W. Zhao, L. Bai, Y. Rao, J. Zhou, and J. Lu, “Unipc: A unified predictor-corrector framework for fast sampling of diffusion models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Later among the works it cites.
A. Swerdlow, R. Xu, and B. Zhou, “Street-view image generation from a bird’s-eye view layout,” IEEE Robotics and Automation Letters , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
H. Li, Z. Yang, Z. Qian, G. Zhao, Y. Huang, J. Yu, and L. Liu, “Dualdiff: Dual-branch diffusion model for autonomous driving with semantic fusion,” in 2025 IEEE International Conference on Robotics and Automation (ICRA) , 2025, accepted for publication in ICRA 2025
2025
Closest in time.
W. Zheng, R. Song, X. Guo, C. Zhang, and L. Chen, “Genad: Generative end-to-end autonomous driving,” in European Conference on Computer Vision . Springer, 2025, pp. 87–104
2025
Closest in time.
X. Li, Y. Zhang, and X. Ye, “Drivingdiffusion: Layout-guided multi-view driving scenarios video generation with latent diffusion model,” in European Conference on Computer Vision . Springer, 2025, pp. 469–485
2025
Closest in time.