Fetching the paper…
Reading the bibliography…
Video generation has advanced significantly, evolving from producing unrealistic outputs to generating videos that appear visually convincing and temporally coherent.
M. A. Fischler and R. C. Bolles, “Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography,”
1981
Earlier work this paper cites.
G. Lowe, “Sift-the scale invariant feature transform,”
2004
Earlier work this paper cites.
M. Muja and D. G. Lowe, “Fast approximate nearest neighbors with automatic algorithm configuration.”
2009
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,”
2013
Earlier work this paper cites.
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” in
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in
2015
Earlier work this paper cites.
L. A. Gatys, A. S. Ecker, and M. Bethge, “A neural algorithm of artistic style,”
2015
Earlier work this paper cites.
T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, X. Chen, and X. Chen, “Improved techniques for training gans,” in
2016
Earlier work this paper cites.
A. Van Den Oord, O. Vinyals
2017
Earlier work this paper cites.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “GANs trained by a two time-scale update rule converge to a local nash equilibrium,” in
2017
Earlier work this paper cites.
T. Unterthiner, S. van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly, “FVD: A new metric for video generation,” in
2019
Earlier work this paper cites.
J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “Arcface: Additive angular margin loss for deep face recognition,” in
2019
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
J. Deng, J. Guo, E. Ververas, I. Kotsia, and S. Zafeiriou, “Retinaface: Single-shot multi-level face localisation in the wild,” in
2020
Earlier work this paper cites.
Z. Teed and J. Deng, “Raft: Recurrent all-pairs field transforms for optical flow,” in
2020
Earlier work this paper cites.
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” in
2021
Earlier work this paper cites.
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” in
2021
Earlier work this paper cites.
M. Ding, Z. Yang, W. Hong, W. Zheng, C. Zhou, D. Yin, J. Lin, X. Zou, Z. Shao, H. Yang
2021
Earlier work this paper cites.
P. Esser, R. Rombach, and B. Ommer, “Taming transformers for high-resolution image synthesis,” in
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
M. Ding, W. Zheng, W. Hong, and J. Tang, “Cogview2: Faster and better text-to-image generation via hierarchical transformers,” in
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
W. Peebles and S. Xie, “Scalable diffusion models with transformers,”
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Z. Xie, Z. Zhang, Y. Cao, Y. Lin, J. Bao, Z. Yao, Q. Dai, and H. Hu, “Simmim: A simple framework for masked image modeling,” in
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Huang, L. Sigal, K. M. Yi, O. Wang, and J.-Y. Lee, “Inve: Interactive neural video editing,”
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
D. Ceylan, C.-H. P. Huang, and N. J. Mitra, “Pix2video: Video editing using image diffusion,” in
OpenAI, “Sora,” Accessed February 15, 2024 [Online]
2024
Later among the works it cites.
K. Team, “Kling,” Accessed December 9, 2024 [Online]
2024
Later among the works it cites.
runway, “Gen-3,” Accessed June 17, 2024 [Online]
2024
Later among the works it cites.
T. Team, “Hunyuanvideo: A systematic framework for large video generative models,” 2024
2024
Later among the works it cites.
P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, H. Saini, Y. Levi, D. Lorenz, A. Sauer, F. Boesel
2024
Later among the works it cites.
Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y. Yang, W. Hong, X. Zhang, G. Feng
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Y. He, M. Xia, H. Chen, X. Cun, Y. Gong, J. Xing, Y. Zhang, X. Wang, C. Weng, Y. Shan
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Y. Wang, X. Chen, X. Ma, S. Zhou, Z. Huang, Y. Wang, C. Yang, Y. He, J. Yu, P. Yang
2023
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei
2024
Later among the works it cites.
T. Cheng, L. Song, Y. Ge, W. Liu, X. Wang, and Y. Shan, “Yolo-world: Real-time open-vocabulary object detection,” in
2024
Later among the works it cites.
G. Fang, W. Yan, Y. Guo, J. Han, Z. Jiang, H. Xu, S. Liao, and X. Liang, “Humanrefiner: Benchmarking abnormal human generation and refining with coarse-to-fine pose-reversible guidance,” in
2024
Later among the works it cites.
N. Karaev, I. Rocco, B. Graham, N. Neverova, A. Vedaldi, and C. Rupprecht, “Cotracker: It is better to track together,” in
2024
Later among the works it cites.
2024
Later among the works it cites.
X. Sun, Y. Chen, Y. Huang, R. Xie, J. Zhu, K. Zhang, S. Li, Z. Yang, J. Han, X. Shu
2024
Later among the works it cites.
V. Balazadeh, M. Ataei, H. Cheong, A. Hosein Khasahmadi, and R. G. Krishnan, “Synthetic vision: Training vision-language models to understand physics,”
2024
Later among the works it cites.
H. Al-Tahan, Q. Garrido, R. Balestriero, D. Bouchacourt, C. Hazirbas, and M. Ibrahim, “Unibench: Visual reasoning requires rethinking vision-language beyond scaling,” in
2024
Later among the works it cites.
X. Yang, L. Zhu, H. Fan, and Y. Yang, “Videograin: Modulating space-time attention for multi-grained video editing,” in
2025
Closest in time.
2025
Closest in time.
G. Team, “Veo2,” Accessed December 18, 2024 [Online]
2025
Closest in time.
2025
Closest in time.
W. Team, “Wan: Open and advanced large-scale video generative models,” 2025
2025
Closest in time.
S. Team, 2025. [Online]. Available:
2025
Closest in time.
W. Fan, C. Si, J. Song, Z. Yang, Y. He, L. Zhuo, Z. Huang, Z. Dong, J. He, D. Pan
2025
Closest in time.
C. Si, W. Fan, Z. Lv, Z. Huang, Y. Qiao, and Z. Liu, “Repvideo: Rethinking cross-layer representation for video generation,”
2025
Closest in time.
2025
Closest in time.
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang
2025
Closest in time.
J. Yang, S. Yang, A. W. Gupta, R. Han, L. Fei-Fei, and S. Xie, “Thinking in space: How multimodal large language models see, remember, and recall spaces,” in
2025
Closest in time.