Fetching the paper…
Reading the bibliography…
Generative models offer a scalable and flexible paradigm for simulating complex environments, yet current approaches fall short in addressing the domain-specific requirements of autonomous driving - such as multi-agent interactions, fine-grained control, and multi-camera consistency.
A bi-symmetric log transformation for wide-range data
J. B. W. Webber · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Earlier work this paper cites.
Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network
W. Shi, J. Caballero, F. Huszar, J. Totz, A. P. Aitken, R. Bishop, D. Rueckert, and Z. Wang · 2016
Earlier work this paper cites.
Perceptual losses for real-time style transfer and super-resolution
J. Johnson, A. Alahi, and L. Fei-Fei · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2016
Earlier work this paper cites.
Neural discrete representation learning
A. van den Oord, O. Vinyals, and K. Kavukcuoglu · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Earlier work this paper cites.
Spectral normalization for generative adversarial networks
T. Miyato, T. Kataoka, M. Koyama, and Y. Yoshida · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
T. Unterthiner, S. Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly · 2018
Earlier work this paper cites.
Making convolutional networks shift-invariant again
R. Zhang · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
P. Esser, R. Rombach, and B. Ommer · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
VideoGPT: Video generation using vq-vae and transformers
W. Yan, Y. Zhang, P. Abbeel, and A. Srinivas · 2021
Earlier work this paper cites.
CCVS: Context-aware controllable video synthesis
G. L. Moing, J. Ponce, and C. Schmid · 2021
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Earlier work this paper cites.
Long video generation with time-agnostic vqgan and time-sensitive transformer
S. Ge, T. Hayes, H. Yang, X. Yin, G. Pang, D. Jacobs, J.-B. Huang, and D. Parikh · 2022
Cited alongside, same era.
HARP: Autoregressive latent video prediction with high-fidelity image generator
Y. Seo, K. Lee, F. Liu, S. James, and P. Abbeel · 2022
Cited alongside, same era.
General-purpose, long-context autoregressive modeling with Perceiver AR
C. Hawthorne, A. Jaegle, C. Cangea, S. Borgeaud, C. Nash, M. Malinowski, S. Dieleman, O. Vinyals, M. Botvinick, I. Simon, H. Sheahan, N. Zeghidour, J.-B. Alayrac, J. Carreira, and J. Engel · 2022
Cited alongside, same era.
Gaia-1: A generative world model for autonomous driving
A. Hu, L. Russell, H. Yeo, Z. Murez, G. Fedoseev, A. Kendall, J. Shotton, and G. Corrado · 2023
Cited alongside, same era.
Flow matching for generative modeling
Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le · 2023
Cited alongside, same era.
Drivedreamer4d: World models are effective data machines for 4d driving scene representation
G. Zhao, C. Ni, X. Wang, Z. Zhu, X. Zhang, Y. Wang, G. Huang, X. Chen, B. Wang, Y. Zhang, et al · 2024
Later among the works it cites.
Ltx-video: Realtime video latent diffusion
Y. HaCohen, N. Chiprut, B. Brazowski, D. Shalem, D. Moshe, E. Richardson, E. Levin, G. Shiran, N. Zabari, O. Gordon, P. Panet, S. Weissbuch, V. Kulikov, Y. Bitterman, Z. Melumian, and O. Bibi · 2024
Later among the works it cites.
DINOv2: Learning robust visual features without supervision
M. Oquab, T. Darcet, T. Moutakanni, H. V. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, R. Howes, P.-Y. Huang, H. Xu, V. Sharma, S.-W. Li, W. Galuba, M. Rabbat, M. Assran, N. Ballas, G. Synnaeve, I. Misra, H. Jegou, J. Mairal, P. Labatut, A. Joulin, and P. Bojanowski · 2024
Later among the works it cites.
Movie gen: A cast of media foundation models
A. Polyak, A. Zohar, A. Brown, A. Tjandra, A. Sinha, A. Lee, A. Vyas, B. Shi, C.-Y. Ma, C.-Y. Chuang, et al · 2024
Later among the works it cites.
Fréchet video motion distance: A metric for evaluating motion consistency in videos
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scalable diffusion models with transformers
W. Peebles and S. Xie · 2023
Cited alongside, same era.
Scaling vision transformers to 22 billion parameters
M. Dehghani, J. Djolonga, B. Mustafa, P. Padlewski, J. Heek, J. Gilmer, A. Steiner, M. Caron, R. Geirhos, I. Alabdulmohsin, R. Jenatton, L. Beyer, M. Tschannen, A. Arnab, X. Wang, C. Riquelme, M. Minderer, J. Puigcerver, U. Evci, M. Kumar, S. van Steenkiste, G. F. Elsayed, A. Mahendran, F. Yu, A. Oliver, F. Huot, J. Bastings, M. P. Collier, A. Gritsenko, V. Birodkar, C. Vasconcelos, Y. Tay, T. Mensink, A. Kolesnikov, F. Pavetić, D. Tran, T. Kipf, M. Lučić, X. Zhai, D. Keysers, J. Harmsen, and N. Houlsby · 2023
Cited alongside, same era.
Exploring object-centric temporal modeling for efficient multi-view 3d object detection
S. Wang, Y. Liu, T. Wang, Y. Li, and X. Zhang · 2023
Cited alongside, same era.
Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion models
G. Stein, J. C. Cresswell, R. Hosseinzadeh, Y. Sui, B. L. Ross, V. Villecroze, Z. Liu, A. L. Caterini, J. E. T. Taylor, and G. Loaiza-Ganem · 2023
Cited alongside, same era.
Oneformer: One transformer to rule universal image segmentation
J. Jain, J. Li, M. Chiu, A. Hassani, N. Orlov, and H. Shi · 2023
Cited alongside, same era.
Transformers are sample-efficient world models
V. Micheli, E. Alonso, and F. Fleuret · 2023
Cited alongside, same era.
Temporally consistent transformers for video generation
W. Yan, D. Hafner, S. James, and P. Abbeel · 2023
Cited alongside, same era.
J. Liu, Y. Qu, Q. Yan, X. Zeng, L. Wang, and R. Liao · 2024
Later among the works it cites.
Cosmos tokenizer: A suite of image and video neural tokenizers
NVIDIA · 2024
Later among the works it cites.
Diffusion for world modeling: Visual details matter in atari
E. Alonso, A. Jelley, V. Micheli, A. Kanervisto, A. Storkey, T. Pearce, and F. Fleuret · 2024
Later among the works it cites.
Video generation models as world simulators
OpenAI · 2024
Later among the works it cites.
Introducing gen-3 alpha: A new frontier for video generation
Runway · 2024
Later among the works it cites.
R. Chen, Z. Wu, Y. Liu, Y. Guo, J. Ni, H. Xia, and S. Xia · 2024
Later among the works it cites.
Unleashing generalization of end-to-end autonomous driving with controllable long video generation
E. Ma, L. Zhou, T. Tang, Z. Zhang, D. Han, J. Jiang, K. Zhan, P. Jia, X. Lang, H. Sun, et al · 2024
Later among the works it cites.
M. Hassan, S. Stapf, A. Rahimi, P. Rezende, Y. Haghighi, D. Brüggemann, I. Katircioglu, L. Zhang, X. Chen, S. Saha, et al · 2024
Later among the works it cites.
Dreamdrive: Generative 4d scene modeling from street view images
J. Mao, B. Li, B. Ivanovic, Y. Chen, Y. Wang, Y. You, C. Xiao, D. Xu, M. Pavone, and Y. Wang · 2024
Later among the works it cites.
Cosmos world foundation model platform for physical ai
N. Agarwal, A. Ali, M. Bala, Y. Balaji, E. Barker, T. Cai, P. Chattopadhyay, Y. Chen, Y. Cui, Y. Ding, et al · 2025
Closest in time.
Deep compression autoencoder for efficient high-resolution diffusion models
J. Chen, H. Cai, J. Chen, E. Xie, S. Yang, H. Tang, M. Li, Y. Lu, and S. Han · 2025
Closest in time.
Maskgwm: A generalizable driving world model with video mask reconstruction
J. Ni, Y. Guo, Y. Liu, R. Chen, L. Lu, and Z. Wu · 2025
Closest in time.