Fetching the paper…
Reading the bibliography…
The burgeoning field of Artificial Intelligence Generated Content (AIGC) is witnessing rapid advancements, particularly in video generation.
Image quality assessment: from error visibility to structural similarity
Wang, Z., Bovik, A.C., Sheikh, H.R., Simoncelli, E.P., 2004 · 2004
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation, in: Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18, Springer. pp. 234–241
Ronneberger, O., Fischer, P., Brox, T., 2015 · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics, in: International conference on machine learning, PMLR. pp. 2256–2265
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., Ganguli, S., 2015 · 2015
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Song, Y., Ermon, S., 2019 · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., Abbeel, P., 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P.J., 2020 · 2020
Earlier work this paper cites.
Raft: Recurrent all-pairs field transforms for optical flow, in: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16, Springer. pp. 402–419
Teed, Z., Deng, J., 2020 · 2020
Earlier work this paper cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 1728–1738
Bain, M., Nagrani, A., Varol, G., Zisserman, A., 2021 · 2021
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Dhariwal, P., Nichol, A., 2021 · 2021
Earlier work this paper cites.
Improved denoising diffusion probabilistic models, in: International Conference on Machine Learning, PMLR. pp. 8162–8171
Nichol, A.Q., Dhariwal, P., 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, in: International conference on machine learning, PMLR. pp. 8748–8763
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al., 2021 · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation, in: International Conference on Machine Learning, PMLR. pp. 8821–8831
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., Sutskever, I., 2021 · 2021
Earlier work this paper cites.
Latent video diffusion models for high-fidelity video generation with arbitrary lengths
He, Y., Yang, T., Zhang, Y., Shan, Y., Chen, Q., 2022 · 2022
Cited alongside, same era.
Imagen video: High definition video generation with diffusion models
Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D.P., Poole, B., Norouzi, M., Fleet, D.J., et al., 2022 · 2022
Cited alongside, same era.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Hong, W., Ding, M., Zheng, W., Liu, X., Tang, J., 2022 · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10684–10695
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B., 2022 · 2022
Cited alongside, same era.
A benchmark for controllable text-image-to-video generation
Hu, Y., Luo, C., Chen, Z., 2023 · 2023
Later among the works it cites.
Dreampose: Fashion image-to-video synthesis via stable diffusion
Karras, J., Holynski, A., Wang, T.C., Kemelmacher-Shlizerman, I., 2023 · 2023
Later among the works it cites.
Text2video-zero: Text-to-image diffusion models are zero-shot video generators
Khachatryan, L., Movsisyan, A., Tadevosyan, V., Henschel, R., Wang, Z., Navasardyan, S., Shi, H., 2023 · 2023
Later among the works it cites.
Li, Z., Tucker, R., Snavely, N., Holynski, A., 2023 · 2023
Later among the works it cites.
Videofusion: Decomposed diffusion models for high-quality video generation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10209–10218
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., Schramowski, P., Kundurthy, S., Crowson, K., Schmidt, L., Kaczmarczyk, R., Jitsev, J., 2022 · 2022
Cited alongside, same era.
Make-a-video: Text-to-video generation without text-video data
Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., et al., 2022 · 2022
Cited alongside, same era.
Videocrafter1: Open diffusion models for high-quality video generation
Chen, H., Xia, M., He, Y., Zhang, Y., Cun, X., Yang, S., Xing, J., Liu, Y., Chen, Q., Wang, X., et al., 2023 · 2023
Cited alongside, same era.
Structure and content-guided video synthesis with diffusion models, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 7346–7356
Esser, P., Chiu, J., Atighehchian, P., Granskog, J., Germanidis, A., 2023 · 2023
Cited alongside, same era.
Preserve your own correlation: A noise prior for video diffusion models, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 22930–22941
Ge, S., Nah, S., Liu, G., Poon, T., Tao, A., Catanzaro, B., Jacobs, D., Huang, J.B., Liu, M.Y., Balaji, Y., 2023 · 2023
Cited alongside, same era.
Emu video: Factorizing text-to-video generation by explicit image conditioning
Girdhar, R., Singh, M., Brown, A., Duval, Q., Azadi, S., Rambhatla, S.S., Shah, A., Yin, X., Parikh, D., Misra, I., 2023 · 2023
Cited alongside, same era.
Seer: Language instructed video prediction with latent diffusion models
Gu, X., Wen, C., Song, J., Gao, Y., 2023 · 2023
Cited alongside, same era.
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Guo, Y., Yang, C., Rao, A., Wang, Y., Qiao, Y., Lin, D., Dai, B., 2023 · 2023
Cited alongside, same era.
Luo, Z., Chen, D., Zhang, Y., Huang, Y., Wang, L., Shen, Y., Zhao, D., Zhou, J., Tan, T., 2023 · 2023
Later among the works it cites.
Conditional image-to-video generation with latent flow diffusion models, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 18444–18455
Ni, H., Shi, C., Li, K., Huang, S.X., Min, M.R., 2023 · 2023
Later among the works it cites.
OpenAI, 2023 · 2023
Later among the works it cites.
Pika lab discord server
Pika, I., 2023 · 2023
Later among the works it cites.
Generative multimodal models are in-context learners
Sun, Q., Cui, Y., Zhang, X., Zhang, F., Yu, Q., Luo, Z., Wang, Y., Rao, Y., Liu, J., Huang, T., Wang, X., 2023 · 2023
Later among the works it cites.
Dragnuwa: Fine-grained control in video generation by integrating text, image, and trajectory
Yin, S., Wu, C., Liang, J., Shi, J., Li, H., Ming, G., Duan, N., 2023 · 2023
Later among the works it cites.
MAGVIT: Masked generative video transformer, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Yu, L., Cheng, Y., Sohn, K., Lezama, J., Zhang, H., Chang, H., Hauptmann, A.G., Yang, M.H., Hao, Y., Essa, I., Jiang, L., 2023 · 2023
Later among the works it cites.
Evaluatology: The Science and Engineering of Evaluation
Zhan, J., Wang, L., Gao, W., Li, H., Huang, Y., Wang, C., Li, Y., Yang, Z., Kang, G., Luo, C., Ye, H., Dai, S., Zhang, Z., 2024 · 2024
Closest in time.