Fetching the paper…
Reading the bibliography…
Autonomous driving progress relies on large-scale annotated datasets.
Auto-encoding variational bayes
Kingma, D. P.; and Welling, M. 2013 · 2013
Earlier work this paper cites.
A large-scale car dataset for fine-grained categorization and verification
Yang, L.; Luo, P.; Change Loy, C.; and Tang, X. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; and Hochreiter, S. 2017 · 2017
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Unterthiner, T.; Van Steenkiste, S.; Kurach, K.; Marinier, R.; Michalski, M.; and Gelly, S. 2018 · 2018
Earlier work this paper cites.
nuscenes: A multimodal dataset for autonomous driving
Caesar, H.; Bankiti, V.; Lang, A. H.; Vora, S.; Liong, V. E.; Xu, Q.; Krishnan, A.; Pan, Y.; Baldan, G.; and Beijbom, O. 2020 · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Earlier work this paper cites.
Denoising Diffusion Implicit Models
Song, J.; Meng, C.; and Ermon, S. 2020 · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
CC-3DT: Panoramic 3D object tracking via cross-camera fusion
Fischer, T.; Yang, Y.-H.; Kumar, S.; Sun, M.; and Yu, F. 2022 · 2022
Earlier work this paper cites.
An image is worth one word: Personalizing text-to-image generation using textual inversion
Gal, R.; Alaluf, Y.; Atzmon, Y.; Patashnik, O.; Bermano, A. H.; Chechik, G.; and Cohen-Or, D. 2022 · 2022
Earlier work this paper cites.
Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers
Li, Z.; Wang, W.; Li, H.; Xie, E.; Sima, C.; Lu, T.; Qiao, Y.; and Dai, J. 2022 · 2022
Earlier work this paper cites.
Petr: Position embedding transformation for multi-view 3d object detection
Liu, Y.; Wang, T.; Zhang, X.; and Sun, J. 2022 · 2022
Earlier work this paper cites.
TripletTrack: 3D object tracking using triplet embeddings and LSTM
Marinello, N.; Proesmans, M.; and Van Gool, L. 2022 · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Cited alongside, same era.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Cited alongside, same era.
Synthetic data from diffusion models improves imagenet classification
Azizi, S.; Kornblith, S.; Saharia, C.; Norouzi, M.; and Fleet, D. J. 2023 · 2023
Cited alongside, same era.
Custom-edit: Text-guided image editing with customized diffusion models
Choi, J.; Choi, Y.; Kim, Y.; Kim, J.; and Yoon, S. 2023 · 2023
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Ruiz, N.; Li, Y.; Jampani, V.; Pritch, Y.; Rubinstein, M.; and Aberman, K. 2023 · 2023
Later among the works it cites.
Fake it till you make it: Learning transferable representations from synthetic imagenet clones
Sarıyıldız, M. B.; Alahari, K.; Larlus, D.; and Kalantidis, Y. 2023 · 2023
Later among the works it cites.
Street-View Image Generation from a Bird’s-Eye View Layout
Swerdlow, A.; Xu, R.; and Zhou, B. 2023 · 2023
Later among the works it cites.
Panacea: Panoramic and Controllable Video Generation for Autonomous Driving
Wen, Y.; Zhao, Y.; Liu, Y.; Jia, F.; Wang, Y.; Luo, C.; Zhang, C.; Wang, T.; Sun, X.; and Zhang, X. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Magicdrive: Street view generation with diverse 3d geometry control
Gao, R.; Chen, K.; Xie, E.; Hong, L.; Li, Z.; Yeung, D.-Y.; and Xu, Q. 2023 · 2023
Cited alongside, same era.
SVDiff: Compact Parameter Space for Diffusion Fine-Tuning
Han, L.; Li, Y.; Zhang, H.; Milanfar, P.; Metaxas, D. N.; and Yang, F. 2023 · 2023
Cited alongside, same era.
VideoBooth: Diffusion-based Video Generation with Image Prompts
Jiang, Y.; Wu, T.; Yang, S.; Si, C.; Lin, D.; Qiao, Y.; Loy, C. C.; and Liu, Z. 2023 · 2023
Cited alongside, same era.
Multi-Concept Customization of Text-to-Image Diffusion
Kumari, N.; Zhang, B.; Zhang, R.; Shechtman, E.; and Zhu, J. 2023 · 2023
Cited alongside, same era.
Li, X.; Zhang, Y.; and Ye, X. 2023 · 2023
Cited alongside, same era.
WoVoGen: World volume-aware diffusion for controllable multi-camera driving scene generation
Lu, J.; Huang, Z.; Zhang, J.; Yang, Z.; and Zhang, L. 2023 · 2023
Cited alongside, same era.
Subject-diffusion: Open domain personalized text-to-image generation without test-time fine-tuning
Ma, J.; Liang, J.; Chen, C.; and Lu, H. 2023 · 2023
Cited alongside, same era.
Yang, K.; Ma, E.; Peng, J.; Guo, Q.; Lin, D.; and Yu, K. 2023 · 2023
Later among the works it cites.
Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models
Ye, H.; Zhang, J.; Liu, S.; Han, X.; and Yang, W. 2023 · 2023
Later among the works it cites.
Magvit: Masked generative video transformer
Yu, L.; Cheng, Y.; Sohn, K.; Lezama, J.; Zhang, H.; Chang, H.; Hauptmann, A. G.; Yang, M.-H.; Hao, Y.; Essa, I.; et al. 2023 · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
Zhang, L.; Rao, A.; and Agrawala, M. 2023 · 2023
Later among the works it cites.
Video generation models as world simulators
Brooks, T.; Peebles, B.; Holmes, C.; DePue, W.; Guo, Y.; Jing, L.; Schnurr, D.; Taylor, J.; Luhman, T.; Luhman, E.; Ng, C.; Wang, R.; and Ramesh, A. 2024 · 2024
Closest in time.
Scaling laws of synthetic images for model training… for now
Fan, L.; Chen, K.; Krishnan, D.; Katabi, D.; Isola, P.; and Tian, Y. 2024 · 2024
Closest in time.
Dataset diffusion: Diffusion-based synthetic data generation for pixel-level semantic segmentation
Nguyen, Q.; Vu, T.; Tran, A.; and Nguyen, K. 2024 · 2024
Closest in time.
Fastcomposer: Tuning-free multi-subject image generation with localized attention
Xiao, G.; Yin, T.; Freeman, W. T.; Durand, F.; and Han, S. 2024 · 2024
Closest in time.