Fetching the paper…
Reading the bibliography…
Video generation has witnessed significant advancements, yet evaluating these models remains a challenge.
T. B. Fitzpatrick, “The validity and practicality of sun-reactive skin types i through vi,”
1988
Earlier work this paper cites.
2012
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,”
2013
Earlier work this paper cites.
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. C. Courville, and Y. Bengio, “Generative adversarial nets,” in
2014
Earlier work this paper cites.
M. Mirza and S. Osindero, “Conditional generative adversarial nets,”
2014
Earlier work this paper cites.
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” in
2015
Earlier work this paper cites.
T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, X. Chen, and X. Chen, “Improved techniques for training gans,” in
2016
Earlier work this paper cites.
J. Xu, T. Mei, T. Yao, and Y. Rui, “Msr-vtt: A large video description dataset for bridging video and language,” in
2016
Earlier work this paper cites.
A. Van Den Oord, O. Vinyals
2017
Earlier work this paper cites.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “GANs trained by a two time-scale update rule converge to a local nash equilibrium,” in
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
T. Karras, T. Aila, S. Laine, and J. Lehtinen, “Progressive growing of GANs for improved quality, stability, and variation,” in
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
T. Karras, S. Laine, and T. Aila, “A style-based generator architecture for generative adversarial networks,” in
2019
Earlier work this paper cites.
——, “FVD: A new metric for video generation,” in
2019
Earlier work this paper cites.
D. Li, T. Jiang, and M. Jiang, “Quality assessment of in-the-wild videos,” in
2019
Earlier work this paper cites.
T. Karras, S. Laine, M. Aittala, J. Hellsten, J. Lehtinen, and T. Aila, “Analyzing and improving the image quality of StyleGAN,” in
2020
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in
2020
Earlier work this paper cites.
Z. Teed and J. Deng, “Raft: Recurrent all-pairs field transforms for optical flow,” in
2020
Earlier work this paper cites.
Y. Fang, H. Zhu, Y. Zeng, K. Ma, and Z. Wang, “Perceptual quality assessment of smartphone photography,” in
2020
Earlier work this paper cites.
J. Deng, J. Guo, E. Ververas, I. Kotsia, and S. Zafeiriou, “Retinaface: Single-shot multi-level face localisation in the wild,” in
2020
Earlier work this paper cites.
S. P. Huntington, “The clash of civilizations?” in
2020
Earlier work this paper cites.
T. Karras, M. Aittala, S. Laine, E. Härkönen, J. Hellsten, J. Lehtinen, and T. Aila, “Alias-free generative adversarial networks,” in
2021
Earlier work this paper cites.
Y. Jiang, Z. Huang, X. Pan, C. C. Loy, and Z. Liu, “Talk-to-Edit: Fine-grained facial editing via dialog,” in
2021
Earlier work this paper cites.
P. Esser, R. Rombach, and B. Ommer, “Taming transformers for high-resolution image synthesis,” in
2021
Earlier work this paper cites.
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” in
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark
2021
Earlier work this paper cites.
Z. Tu, Y. Wang, N. Birkbeck, B. Adsumilli, and A. C. Bovik, “Ugc-vqa: Benchmarking blind video quality assessment for user generated content,”
2021
Earlier work this paper cites.
P. Dhariwal and A. Nichol, “Diffusion models beat GANs on image synthesis,” in
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging properties in self-supervised vision transformers,” in
2021
Earlier work this paper cites.
J. Ke, Q. Wang, Y. Wang, P. Milanfar, and F. Yang, “MUSIQ: multi-scale image quality transformer,”
2021
Earlier work this paper cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” 2021
2021
Earlier work this paper cites.
M. Bain, A. Nagrani, G. Varol, and A. Zisserman, “Frozen in time: A joint video and image encoder for end-to-end retrieval,” in
2021
Earlier work this paper cites.
J. Fu, S. Li, Y. Jiang, K.-Y. Lin, C. Qian, C. C. Loy, W. Wu, and Z. Liu, “Stylegan-human: A data-centric odyssey of human generation,” in
2022
Earlier work this paper cites.
Y. Jiang, S. Yang, H. Qju, W. Wu, C. C. Loy, and Z. Liu, “Text2human: Text-driven controllable human image generation,”
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
R. Villegas, M. Babaeizadeh, P.-J. Kindermans, H. Moraldo, H. Zhang, M. T. Saffar, S. Castro, J. Kunze, and D. Erhan, “Phenaki: Variable length video generation from open domain textual descriptions,” in
2022
Earlier work this paper cites.
S. Ge, T. Hayes, H. Yang, X. Yin, G. Pang, D. Jacobs, J.-B. Huang, and D. Parikh, “Long video generation with time-agnostic vqgan and time-sensitive transformer,” in
2022
Earlier work this paper cites.
M. Ding, W. Zheng, W. Hong, and J. Tang, “Cogview2: Faster and better text-to-image generation via hierarchical transformers,” in
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
H. Wu, C. Chen, J. Hou, L. Liao, A. Wang, W. Sun, Q. Yan, and W. Lin, “Fast-vqa: Efficient end-to-end video quality assessment with fragment sampling,” in
2022
Earlier work this paper cites.
S. Gu, D. Chen, J. Bao, F. Wen, B. Zhang, D. Chen, L. Yuan, and B. Guo, “Vector quantized diffusion model for text-to-image synthesis,” in
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet, “Video diffusion models,”
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
OpenAI, “Dalle 3 system card,”
2022
Earlier work this paper cites.
P. Schramowski, C. Tauchmann, and K. Kersting, “Can machines help us answering question 16 in datasheets, and in turn reflecting on inappropriate content?” in
2022
Earlier work this paper cites.
LAION-AI, “aesthetic-predictor,”
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Cited alongside, same era.
J. Fu, S. Li, Y. Jiang, K.-Y. Lin, W. Wu, and Z. Liu, “Unitedhuman: Harnessing multi-source data for high-resolution human generation,” in
2023
Cited alongside, same era.
Y. Jiang, Z. Huang, T. Wu, X. Pan, C. C. Loy, and Z. Liu, “Talk-to-edit: Fine-grained 2d and 3d facial editing via dialog,”
2023
Cited alongside, same era.
Z. Luo, D. Chen, Y. Zhang, Y. Huang, L. Wang, Y. Shen, D. Zhao, J. Zhou, and T. Tan, “VideoFusion: Decomposed diffusion models for high-quality video generation,” in
2023
Cited alongside, same era.
A. Blattmann, R. Rombach, H. Ling, T. Dockhorn, S. W. Kim, S. Fidler, and K. Kreis, “Align your latents: High-resolution video synthesis with latent diffusion models,” in
L. Struppek, D. Hintersdorf, F. Friedrich, P. Schramowski, K. Kersting
2023
Later among the works it cites.
2023
Later among the works it cites.
A. Basu, R. V. Babu, and D. Pruthi, “Inspecting the geographical representativeness of images from text-to-image models,” in
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Y. Wang, X. Chen, X. Ma, S. Zhou, Z. Huang, Y. Wang, C. Yang, Y. He, J. Yu, P. Yang
2023
Cited alongside, same era.
2023
Cited alongside, same era.
L. Yu, Y. Cheng, K. Sohn, J. Lezama, H. Zhang, H. Chang, A. G. Hauptmann, M.-H. Yang, Y. Hao, I. Essa
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Later among the works it cites.
J. Cho, A. Zala, and M. Bansal, “Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models,” in
2023
Later among the works it cites.
notAI tech, “Nudenet,”
2023
Later among the works it cites.
Z. Li, Z.-L. Zhu, L.-H. Han, Q. Hou, C.-L. Guo, and M.-M. Cheng, “Amt: All-pairs multi-field transforms for efficient frame interpolation,” in
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in
2023
Later among the works it cites.
LAION-AI, “Nsfw detector,”
2023
Later among the works it cites.
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing
2023
Later among the works it cites.
OpenAI, “GPT-4 technical report,”
2023
Later among the works it cites.
“Gen-2,” Accessed September 25, 2023 [Online]
2023
Later among the works it cites.
“Pika 1.0,” Accessed December 28, 2023 [Online]
2023
Later among the works it cites.
2023
Later among the works it cites.
O. Bar-Tal, H. Chefer, O. Tov, C. Herrmann, R. Paiss, S. Zada, A. Ephrat, J. Hur, Y. Li, T. Michaeli
2024
Closest in time.
2024
Closest in time.
W. Wang, J. Liu, Z. Lin, J. Yan, S. Chen, C. Low, T. Hoang, J. Wu, J. H. Liew, H. Yan
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
W. Weng, R. Feng, Y. Wang, Q. Dai, C. Wang, D. Yin, Z. Zhao, K. Qiu, J. Bao, Y. Yuan
2024
Closest in time.
Y. Zeng, G. Wei, J. Zheng, J. Zou, Y. Wei, Y. Zhang, and H. Li, “Make pixels dance: High-dynamic video generation,” in
2024
Closest in time.
2024
Closest in time.
X. Shi, Z. Huang, F.-Y. Wang, W. Bian, D. Li, Y. Zhang, M. Zhang, K. C. Cheung, S. See, H. Qin
2024
Closest in time.
2024
Closest in time.
Z. Huang, Y. He, J. Yu, F. Zhang, C. Si, Y. Jiang, Y. Zhang, T. Wu, Q. Jin, N. Chanpaisit, Y. Wang, X. Chen, L. Wang, D. Lin, Y. Qiao, and Z. Liu, “VBench: Comprehensive benchmark suite for video generative models,” in
2024
Closest in time.
Y. Guo, C. Yang, A. Rao, Y. Wang, Y. Qiao, D. Lin, and B. Dai, “Animatediff: Animate your personalized text-to-image diffusion models without specific tuning,” in
2024
Closest in time.
2024
Closest in time.
H. Ouyang, Q. Wang, Y. Xiao, Q. Bai, J. Zhang, K. Zheng, X. Zhou, Q. Chen, and Y. Shen, “Codef: Content deformation fields for temporally consistent video processing,” in
2024
Closest in time.
M. Ku, T. Li, K. Zhang, Y. Lu, X. Fu, W. Zhuang, and W. Chen, “Imagenhub: Standardizing the evaluation of conditional image generation models,” in
2024
Closest in time.
Y. Liu, X. Cun, X. Liu, X. Wang, Y. Zhang, H. Chen, Y. Liu, T. Zeng, R. Chan, and Y. Shan, “Evalcrafter: Benchmarking and evaluating large video generation models,” in
2024
Closest in time.
F. Fan, C. Luo, W. Gao, and J. Zhan, “Aigcbench: Comprehensive evaluation of image-to-video content generated by ai,”
2024
Closest in time.
Y. Zhang, Z. Xing, Y. Zeng, Y. Fang, and K. Chen, “Pia: Your personalized image animator via plug-and-play modules in text-to-image models,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
T. Lee, M. Yasunaga, C. Meng, Y. Mai, J. S. Park, A. Gupta, Y. Zhang, D. Narayanan, H. Teufel, M. Bellagente
2024
Closest in time.
2024
Closest in time.
X. Shen, C. Du, T. Pang, M. Lin, Y. Wong, and M. Kankanhalli, “Finetuning text-to-image diffusion models for fairness,” in
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
“Pexels, royalty-free stock footage website,”
2024
Closest in time.
“Pixabay, royalty-free stock footage website,”
2024
Closest in time.
“Kling,” Accessed June 6, 2024 [Online]
2024
Closest in time.
“Gen-3,” Accessed June 17, 2024 [Online]
2024
Closest in time.