Fetching the paper…
Reading the bibliography…
Research on video generation has recently made tremendous progress, enabling high-quality videos to be generated from text prompts or images.
Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography
M. A. Fischler and R. C. Bolles · 1981
Earlier work this paper cites.
A comparative analysis of ransac techniques leading to adaptive real-time random sample consensus
R. Raguram, J.-M. Frahm, and M. Pollefeys · 2008
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Demystifying mmd gans
M. Bińkowski, D. J. Sutherland, M. Arbel, and A. Gretton · 2018
Earlier work this paper cites.
Stereo magnification: Learning view synthesis using multiplane images
T. Zhou, R. Tucker, J. Flynn, G. Fyffe, and N. Snavely · 2018
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
High-fidelity performance metrics for generative models in pytorch, 2020
A. Obukhov, M. Seitzer, P.-W. Wu, S. Zhydenko, J. Kyl, and E. Y.-J. Lin · 2020
Earlier work this paper cites.
SuperGlue: Learning feature matching with graph neural networks
P.-E. Sarlin, D. DeTone, T. Malisiewicz, and A. Rabinovich · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2020
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole · 2020
Earlier work this paper cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
M. Bain, A. Nagrani, G. Varol, and A. Zisserman · 2021
Earlier work this paper cites.
Infinite nature: Perpetual view generation of natural scenes from a single image
A. Liu, R. Tucker, V. Jampani, A. Makadia, N. Snavely, and A. Kanazawa · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
Latent video diffusion models for high-fidelity long video generation
Y. He, T. Yang, Y. Zhang, Y. Shan, and Q. Chen · 2022
Earlier work this paper cites.
Imagen video: High definition video generation with diffusion models
J. Ho, W. Chan, C. Saharia, J. Whang, R. Gao, A. Gritsenko, D. P. Kingma, B. Poole, M. Norouzi, D. J. Fleet, and T. Salimans · 2022
Earlier work this paper cites.
Video diffusion models
J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet · 2022
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2022
Earlier work this paper cites.
Elucidating the design space of diffusion-based generative models
T. Karras, M. Aittala, T. Aila, and S. Laine · 2022
Earlier work this paper cites.
Infinitenature-zero: Learning perpetual view generation of natural scenes from single images
Z. Li, Q. Wang, N. Snavely, and A. Kanazawa · 2022
Earlier work this paper cites.
Hierarchical text-conditional image generation with clip latents
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. Denton, S. K. S. Ghasemipour, B. K. Ayan, S. S. Mahdavi, R. G. Lopes, T. Salimans, J. Ho, D. J. Fleet, and M. Norouzi · 2022
Cited alongside, same era.
Make-a-video: Text-to-video generation without text-video data
U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni, D. Parikh, S. Gupta, and Y. Taigman · 2022
Cited alongside, same era.
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
J. Z. Wu, Y. Ge, X. Wang, S. W. Lei, Y. Gu, W. Hsu, Y. Shan, X. Qie, and M. Z. Shou · 2022
Cited alongside, same era.
Videogen: A reference-guided latent diffusion approach for high definition text-to-video generation
X. Li, W. Chu, Y. Wu, W. Yuan, F. Liu, Q. Zhang, F. Li, H. Feng, E. Ding, and J. Wang · 2023
Later among the works it cites.
Zero-1-to-3: Zero-shot one image to 3d object
R. Liu, R. Wu, B. V. Hoorick, P. Tokmakov, S. Zakharov, and C. Vondrick · 2023
Later among the works it cites.
Syncdreamer: Generating multiview-consistent images from a single-view image
Y. Liu, C. Lin, Z. Zeng, X. Long, L. Liu, T. Komura, and W. Wang · 2023
Later among the works it cites.
Wonder3d: Single image to 3d using cross-domain diffusion
X. Long, Y.-C. Guo, C. Lin, Y. Liu, Z. Dou, L. Liu, Y. Ma, S.-H. Zhang, M. Habermann, C. Theobalt, et al · 2023
Later among the works it cites.
State of the art on diffusion models for visual computing
R. Po, W. Yifan, V. Golyanik, K. Aberman, J. T. Barron, A. H. Bermano, E. R. Chan, T. Dekel, A. Holynski, A. Kanazawa, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Universal guidance for diffusion models
A. Bansal, H.-M. Chu, A. Schwarzschild, S. Sengupta, M. Goldblum, J. Geiping, and T. Goldstein · 2023
Cited alongside, same era.
Multidiffusion: Fusing diffusion paths for controlled image generation
O. Bar-Tal, L. Yariv, Y. Lipman, and T. Dekel · 2023
Cited alongside, same era.
Stable video diffusion: Scaling latent video diffusion models to large datasets
A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, D. Lorenz, Y. Levi, Z. English, V. Voleti, A. Letts, V. Jampani, and R. Rombach · 2023
Cited alongside, same era.
Align your latents: High-resolution video synthesis with latent diffusion models
A. Blattmann, R. Rombach, H. Ling, T. Dockhorn, S. W. Kim, S. Fidler, and K. Kreis · 2023
Cited alongside, same era.
Diffdreamer: Towards consistent unsupervised single-view scene extrapolation with conditional diffusion models
S. Cai, E. R. Chan, S. Peng, M. Shahbazi, A. Obukhov, L. Van Gool, and G. Wetzstein · 2023
Cited alongside, same era.
Pix2video: Video editing using image diffusion
D. Ceylan, C.-H. Huang, and N. J. Mitra · 2023
Cited alongside, same era.
GeNVS: Generative novel view synthesis with 3D-aware diffusion models
E. R. Chan, K. Nagano, M. A. Chan, A. W. Bergman, J. J. Park, A. Levy, M. Aittala, S. D. Mello, T. Karras, and G. Wetzstein · 2023
Cited alongside, same era.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
N. Ruiz, Y. Li, V. Jampani, Y. Pritch, M. Rubinstein, and K. Aberman · 2023
Later among the works it cites.
Zero123++: a single image to consistent multi-view diffusion base model
R. Shi, H. Chen, Z. Zhang, M. Liu, C. Xu, X. Wei, L. Chen, C. Zeng, and H. Su · 2023
Later among the works it cites.
Make-it-3d: High-fidelity 3d creation from a single image with diffusion prior
J. Tang, T. Wang, B. Zhang, T. Zhang, R. Yi, L. Ma, and D. Chen · 2023
Later among the works it cites.
Diffusion with forward models: Solving stochastic inverse problems without direct supervision
A. Tewari, T. Yin, G. Cazenavette, S. Rezchikov, J. B. Tenenbaum, F. Durand, W. T. Freeman, and V. Sitzmann · 2023
Later among the works it cites.
Consistent view synthesis with pose-guided diffusion models
H.-Y. Tseng, Q. Li, C. Kim, S. Alsisan, J.-B. Huang, and J. Kopf · 2023
Later among the works it cites.
Videocomposer: Compositional video synthesis with motion controllability
X. Wang, H. Yuan, S. Zhang, D. Chen, J. Wang, Y. Zhang, Y. Shen, D. Zhao, and J. Zhou · 2023
Later among the works it cites.
Motionctrl: A unified and flexible motion controller for video generation
Z. Wang, Z. Yuan, X. Wang, T. Chen, M. Xia, P. Luo, and Y. Shan · 2023
Later among the works it cites.
Dmv3d: Denoising multi-view diffusion using 3d large reconstruction model
Y. Xu, H. Tan, F. Luan, S. Bi, P. Wang, J. Li, Z. Shi, K. Sunkavalli, G. Wetzstein, Z. Xu, and K. Zhang · 2023
Later among the works it cites.
Dragnuwa: Fine-grained control in video generation by integrating text, image, and trajectory
S. Yin, C. Wu, J. Liang, J. Shi, H. Li, G. Ming, and N. Duan · 2023
Later among the works it cites.
Video probabilistic diffusion models in projected latent space
S. Yu, K. Sohn, S. Kim, and J. Shin · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
L. Zhang, A. Rao, and M. Agrawala · 2023
Later among the works it cites.
Diffcollage: Parallel generation of large content with diffusion models
Q. Zhang, J. Song, X. Huang, Y. Chen, and M. yu Liu · 2023
Later among the works it cites.
Motiondirector: Motion customization of text-to-video diffusion models
R. Zhao, Y. Gu, J. Z. Wu, D. J. Zhang, J. Liu, W. Wu, J. Keppo, and M. Z. Shou · 2023
Later among the works it cites.
Video generation models as world simulators
T. Brooks, B. Peebles, C. Holmes, W. DePue, Y. Guo, L. Jing, D. Schnurr, J. Taylor, T. Luhman, E. Luhman, C. Ng, R. Wang, and A. Ramesh · 2024
Closest in time.
Generative rendering: Controllable 4d-guided video generation with 2d diffusion models
S. Cai, D. Ceylan, M. Gadelha, C.-H. Huang, T. Wang, and G. Wetzstein · 2024
Closest in time.
Cameractrl: Enabling camera control for text-to-video generation
H. He, Y. Xu, Y. Guo, G. Wetzstein, B. Dai, H. Li, and C. Yang · 2024
Closest in time.