Fetching the paper…
Reading the bibliography…
Several text-to-video diffusion models have demonstrated commendable capabilities in synthesizing high-quality video content.
Brox, T., Malik, J.: Large displacement optical flow: Descriptor matching in variational motion estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence 33
2010
Earlier work this paper cites.
Sutskever, I., Martens, J., Hinton, G.E.: Generating text with recurrent neural networks. In: Proceedings of the 28th international conference on machine learning (ICML-11). pp. 1017–1024 (2011)
2011
Earlier work this paper cites.
Vondrick, C., Pirsiavash, H., Torralba, A.: Generating videos with scene dynamics. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2016)
2016
Earlier work this paper cites.
Sun, Y., Pan, J., Liu, C., Kang, S.B.: Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2018)
2018
Earlier work this paper cites.
Abdal, R., Qin, Y., Wonka, P.: Image2stylegan: How to embed images into the stylegan latent space? In: 2019 IEEE/CVF International Conference on Computer Vision (ICCV). pp. 4431–4440 (2019). https://doi.org/10.1109/ICCV.2019.00453
2019
Earlier work this paper cites.
Chan, T.Y., Loy, C.C., Tang, X.: Every picture tells a story: Dense video captioning without temporal supervision. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2019)
2019
Earlier work this paper cites.
Zhang, H., Sindagi, V., Patel, V.M.: Image de-raining using a conditional generative adversarial network. IEEE Transactions on Circuits and Systems for Video Technology 30
2019
Earlier work this paper cites.
Zhou, Z., Siddiquee, M.M.R., Tajbakhsh, N., Liang, J.: Unet++: Redesigning skip connections to exploit multiscale features in image segmentation. IEEE Transactions on Medical Imaging 39
2019
Earlier work this paper cites.
Kaplan, J., McCandlish, S., Henighan, T., Brown, T.B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., Amodei, D.: Scaling laws for neural language models (2020)
2020
Earlier work this paper cites.
Karras, T., Laine, S., Aittala, M., Hellsten, J., Lehtinen, J., Aila, T.: Analyzing and improving the image quality of stylegan (2020)
2020
Earlier work this paper cites.
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P.J.: Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research 21
2020
Cited alongside, same era.
Yang, K., Cao, Y., Zhang, Y., Tang, M., Aberg, D., Sadigh, B., Zhou, F.: Self-supervised learning and prediction of microstructure evolution with recurrent neural networks (2020)
2020
Cited alongside, same era.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: Learning transferable visual models from natural language supervision (2021)
2021
Cited alongside, same era.
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., Sutskever, I.: Zero-shot text-to-image generation (2021)
2021
Cited alongside, same era.
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models (2022)
2022
Later among the works it cites.
Roos, H.G.: Layer-adapted meshes for weak boundary layers (2022)
2022
Later among the works it cites.
Datta, S., Ku, A., Ramachandran, D., Anderson, P.: Prompt expansion for adaptive text-to-image generation (2023)
2023
Later among the works it cites.
Kelly, C., Hu, L., Yang, C., Tian, Y., Yang, D., Yang, B., Huang, Z., Li, Z., Zou, Y.: Unifiedvisiongpt: Streamlining vision-oriented ai through generalized multimodal framework (2023)
2023
Later among the works it cites.
Liu, S., Zeng, Z., Ren, T., Li, F., Zhang, H., Yang, J., Li, C., Yang, J., Su, H., Zhu, J., Zhang, L.: Grounding dino: Marrying dino with grounded pre-training for open-set object detection (2023)
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models (2021)
2021
Cited alongside, same era.
Xu, H., Zhang, J., Cai, J., Rezatofighi, H., Tao, D.: Gmflow: Learning optical flow via global matching (2021)
2021
Cited alongside, same era.
Fakhravar, H., Tahami, H.: International co-branding and firms finance performance (2022)
2022
Cited alongside, same era.
Hong, W., Ding, M., Zheng, W., Liu, X., Tang, J.: Cogvideo: Large-scale pretraining for text-to-video generation via transformers (2022)
2022
Cited alongside, same era.
Jain, A., Mildenhall, B., Barron, J.T., Abbeel, P., Poole, B.: Zero-shot text-guided object generation with dream fields. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 857–866 (2022). https://doi.org/10.1109/CVPR52688.2022.00094
2022
Cited alongside, same era.
Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., Chen, M.: Glide: Towards photorealistic image generation and editing with text-guided diffusion models (2022)
2022
Cited alongside, same era.
Wu, T., He, S., Liu, J., Sun, S., Liu, K., Han, Q.L., Tang, Y.: A brief overview of chatgpt: The history, status quo and potential future development. IEEE/CAA Journal of Automatica Sinica 10
2023
Later among the works it cites.
Xing, J., Xia, M., Zhang, Y., Chen, H., Yu, W., Liu, H., Wang, X., Wong, T.T., Shan, Y.: Dynamicrafter: Animating open-domain images with video diffusion priors (2023)
2023
Later among the works it cites.
Zhang, S., Wang, J., Zhang, Y., Zhao, K., Yuan, H., Qin, Z., Wang, X., Zhao, D., Zhou, J.: I2vgen-xl: High-quality image-to-video synthesis via cascaded diffusion models. CoRR (2023)
2023
Later among the works it cites.
Liu, Y., Zhang, K., Li, Y., Yan, Z., Gao, C., Chen, R., Yuan, Z., Huang, Y., Sun, H., Gao, J., He, L., Sun, L.: Sora: A review on background, technology, limitations, and opportunities of large vision models (2024)
2024
Closest in time.