Fetching the paper…
Reading the bibliography…
High-quality driving video generation is crucial for providing training data for autonomous driving models.
Denoising diffusion implicit models
Song, J.; Meng, C.; and Ermon, S. 2020 · 2010
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P.; and Welling, M. 2013 · 2013
Earlier work this paper cites.
Structure-from-motion revisited
Schonberger, J. L.; and Frahm, J.-M. 2016 · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; and Hochreiter, S. 2017 · 2017
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Unterthiner, T.; Van Steenkiste, S.; Kurach, K.; Marinier, R.; Michalski, M.; and Gelly, S. 2018 · 2018
Earlier work this paper cites.
Stereo Magnification: Learning View Synthesis using Multiplane Images
Zhou, T.; Tucker, R.; Flynn, J.; Fyffe, G.; and Snavely, N. 2018 · 2018
Earlier work this paper cites.
nuscenes: A multimodal dataset for autonomous driving
Caesar, H.; Bankiti, V.; Lang, A. H.; Vora, S.; Liong, V. E.; Xu, Q.; Krishnan, A.; Pan, Y.; Baldan, G.; and Beijbom, O. 2020 · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Dhariwal, P.; and Nichol, A. 2021 · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
Ramesh, A.; Pavlov, M.; Goh, G.; Gray, S.; Voss, C.; Radford, A.; Chen, M.; and Sutskever, I. 2021 · 2021
Cited alongside, same era.
LoFTR: Detector-free local feature matching with transformers
Sun, J.; Shen, Z.; Wang, Y.; Bao, H.; and Zhou, X. 2021 · 2021
Cited alongside, same era.
Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers
Li, Z.; Wang, W.; Li, H.; Xie, E.; Sima, C.; Lu, T.; Qiao, Y.; and Dai, J. 2022 · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Cited alongside, same era.
Diffusers: State-of-the-art diffusion models
Von Platen, P.; Patil, S.; Lozhkov, A.; Cuenca, P.; Lambert, N.; Rasul, K.; Davaadorj, M.; and Wolf, T. 2022 · 2022
Cited alongside, same era.
Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Layoutdiffusion: Controllable diffusion model for layout-to-image generation
Zheng, G.; Zhou, X.; Li, X.; Qi, Z.; Shan, Y.; and Li, X. 2023 · 2023
Later among the works it cites.
MagicDrive: Street View Generation with Diverse 3D Geometry Control
Gao, R.; Chen, K.; Xie, E.; Hong, L.; Li, Z.; Yeung, D.-Y.; and Xu, Q. 2024 · 2024
Closest in time.
CameraCtrl: Enabling Camera Control for Text-to-Video Generation
He, H.; Xu, Y.; Guo, Y.; Wetzstein, G.; Dai, B.; Li, H.; and Yang, C. 2024 · 2024
Closest in time.
Training-free Camera Control for Video Generation
Hou, C.; Wei, G.; Zeng, Y.; and Chen, Z. 2024 · 2024
Closest in time.
Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Blattmann, A.; Dockhorn, T.; Kulal, S.; Mendelevitch, D.; Kilian, M.; Lorenz, D.; Levi, Y.; English, Z.; Voleti, V.; Letts, A.; Jampani, V.; and Rombach, R. 2023 · 2023
Cited alongside, same era.
LayoutDiffuse: Adapting Foundational Diffusion Models for Layout-to-Image Generation
Cheng, J.; Liang, X.; Shi, X.; He, T.; Xiao, T.; and Li, M. 2023 · 2023
Cited alongside, same era.
MV-Diffusion: Motion-aware video diffusion model
Deng, Z.; He, X.; Peng, Y.; Zhu, X.; and Cheng, L. 2023 · 2023
Cited alongside, same era.
Adding conditional control to text-to-image diffusion models
Zhang, L.; Rao, A.; and Agrawala, M. 2023 · 2023
Cited alongside, same era.
Analytisch-geometrische Entwicklungen , volume 2
Plücker, J. 1828
Cited in the paper.
Driving into the future: Multiview visual forecasting and planning with world model for autonomous driving
Wang, Y.; He, J.; Fan, L.; Li, H.; Chen, Y.; and Zhang, Z. 2024a
Cited in the paper.
Bevformer v2: Adapting modern image backbones to bird’s-eye-view recognition via perspective supervision
Yang, C.; Chen, Y.; Tian, H.; Tao, C.; Zhu, X.; Zhang, Z.; Huang, G.; Li, H.; Qiao, Y.; Lu, L.; et al. 2023a
Cited in the paper.
Kuang, Z.; Cai, S.; He, H.; Xu, Y.; Li, H.; Guibas, L.; and Wetzstein, G. 2024 · 2024
Closest in time.
Vivid-ZOO: Multi-View Video Generation with Diffusion Model
Li, B.; Zheng, C.; Zhu, W.; Mai, J.; Zhang, B.; Wonka, P.; and Ghanem, B. 2024 · 2024
Closest in time.
Street-view image generation from a bird’s-eye view layout
Swerdlow, A.; Xu, R.; and Zhou, B. 2024 · 2024
Closest in time.
Motionctrl: A unified and flexible motion controller for video generation
Wang, Z.; Yuan, Z.; Wang, X.; Li, Y.; Chen, T.; Xia, M.; Luo, P.; and Shan, Y. 2024b · 2024
Closest in time.
CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation
Xu, D.; Nie, W.; Liu, C.; Liu, S.; Kautz, J.; Wang, Z.; and Vahdat, A. 2024 · 2024
Closest in time.