Fetching the paper…
Reading the bibliography…
With the explosive popularity of AI-generated content (AIGC), video generation has recently received a lot of attention.
Wang X, Xie L, Dong C, et al (2021) Real-ESRGAN: Training real-world blind super-resolution with pure synthetic data. In: International Conference on Computer Vision Workshops, pp 1905–1914
1914
Earlier work this paper cites.
Chen DL, Dolan WB (2011) Collecting highly parallel data for paraphrase evaluation. In: Annual Meeting of the Association for Computational Linguistics
2011
Earlier work this paper cites.
Wah C, Branson S, Welinder P, et al (2011) The Caltech-UCSD birds-200-2011 dataset
2011
Earlier work this paper cites.
Soomro K, Zamir AR, Shah M (2012) UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv
2012
Earlier work this paper cites.
Goodfellow IJ, Pouget-Abadie J, Mirza M, et al (2014) Generative adversarial nets. In: Conference and Workshop on Neural Information Processing Systems
2014
Earlier work this paper cites.
Heilbron FC, Escorcia V, Ghanem B, et al (2015) Activitynet: A large-scale video benchmark for human activity understanding. In: Conference on Computer Vision and Pattern Recognition
2015
Earlier work this paper cites.
Rohrbach A, Rohrbach M, Tandon N, et al (2015) A dataset for movie description. In: Conference on Computer Vision and Pattern Recognition, pp 3202–3212
2015
Earlier work this paper cites.
Mansimov E, Parisotto E, Ba LJ, et al (2016) Generating images from captions with attention. In: International Conference on Learning Representations
2016
Earlier work this paper cites.
Mathieu M, Couprie C, LeCun Y (2016) Deep multi-scale video prediction beyond mean square error. In: International Conference on Learning Representations
2016
Earlier work this paper cites.
Reed SE, Akata Z, Yan X, et al (2016) Generative adversarial text to image synthesis. In: International Conference on Machine Learning
2016
Earlier work this paper cites.
Vondrick C, Pirsiavash H, Torralba A (2016) Generating videos with scene dynamics. In: Conference and Workshop on Neural Information Processing Systems
2016
Earlier work this paper cites.
Xu J, Mei T, Yao T, et al (2016) MSR-VTT: A large video description dataset for bridging video and language. In: Conference on Computer Vision and Pattern Recognition
2016
Earlier work this paper cites.
Joshi BJ, Stewart K, Shapiro D (2017) Bringing impressionism to life with neural style transfer in Come Swim . In: ACM SIGGRAPH Digital Production Symposium
2017
Earlier work this paper cites.
Pan Y, Qiu Z, Yao T, et al (2017) To create what you tell: Generating videos from captions. In: ACM International Conference on Multimedia
2017
Earlier work this paper cites.
Saito M, Matsumoto E, Saito S (2017) Temporal generative adversarial nets with singular value clipping. In: International Conference on Computer Vision
2017
Earlier work this paper cites.
Zhang H, Xu T, Li H (2017) StackGAN: Text to photo-realistic image synthesis with stacked generative adversarial networks. In: International Conference on Computer Vision
2017
Earlier work this paper cites.
Li Y, Min MR, Shen D, et al (2018) Video generation from text. In: AAAI Conference on Artificial Intelligence
2018
Earlier work this paper cites.
Sun D, Yang X, Liu MY, et al (2018) PWC-Net: CNNs for optical flow using pyramid, warping, and cost volume. In: Conference on Computer Vision and Pattern Recognition
2018
Earlier work this paper cites.
Tulyakov S, Liu M, Yang X, et al (2018) MoCoGAN: Decomposing motion and content for video generation. In: Conference on Computer Vision and Pattern Recognition
2018
Earlier work this paper cites.
Unterthiner T, van Steenkiste S, Kurach K, et al (2018) Towards accurate generative models of video: A new metric & challenges. arXiv
2018
Earlier work this paper cites.
Wang TC, Liu MY, Zhu JY, et al (2018) Video-to-video synthesis. arXiv
2018
Earlier work this paper cites.
Xu T, Zhang P, Huang Q, et al (2018) AttnGAN: Fine-grained text to image generation with attentional generative adversarial networks. In: Conference on Computer Vision and Pattern Recognition
2018
Earlier work this paper cites.
Zhang Y, Li K, Li K, et al (2018) Image super-resolution using very deep residual channel attention networks. In: European Conference on Computer Vision, pp 286–301
2018
Earlier work this paper cites.
Zhou L, Xu C, Corso JJ (2018) Towards automatic learning of procedures from web instructional videos. In: AAAI Conference on Artificial Intelligence
2018
Earlier work this paper cites.
Baek Y, Lee B, Han D, et al (2019) Character region awareness for text detection. In: Conference on Computer Vision and Pattern Recognition
2019
Earlier work this paper cites.
Habibian A, van Rozendaal T, Tomczak JM, et al (2019) Video compression with rate-distortion autoencoders. In: International Conference on Computer Vision
2019
Earlier work this paper cites.
Miech A, Zhukov D, Alayrac JB, et al (2019) HowTo100M: Learning a text-video embedding by watching hundred million narrated video clips. In: International Conference on Computer Vision
2019
Earlier work this paper cites.
Wang X, Wu J, Chen J, et al (2019) Vatex: A large-scale, high-quality multilingual dataset for video-and-language research. In: International Conference on Computer Vision
2019
Cited alongside, same era.
Ho J, Jain A, Abbeel P (2020) Denoising diffusion probabilistic models. In: Conference and Workshop on Neural Information Processing Systems
2020
Cited alongside, same era.
Pessoa J, Aidos H, Tomás P, et al (2020) End-to-end learning of video compression using spatio-temporal autoencoders. In: IEEE Workshop on Signal Processing Systems
2020
Cited alongside, same era.
Zeng Y, Fu J, Chao H (2020) Learning joint spatial-temporal transformations for video inpainting. In: European Conference on Computer Vision
2020
Cited alongside, same era.
Bain M, Nagrani A, Varol G, et al (2021) Frozen in time: A joint video and image encoder for end-to-end retrieval. In: International Conference on Computer Vision
Skorokhodov I, Tulyakov S, Elhoseiny M (2022) StyleGAN-V: A continuous video generator with the price, image quality and perks of stylegan2. In: Conference on Computer Vision and Pattern Recognition, pp 3626–3636
2022
Later among the works it cites.
Villegas R, Babaeizadeh M, Kindermans P, et al (2022) Phenaki: Variable length video generation from open domain textual description. arXiv
2022
Later among the works it cites.
Wang P, Wang X, Wang F, et al (2022) KVT: k-nn attention for boosting vision transformers. In: European Conference on Computer Vision
2022
Later among the works it cites.
Wu C, Liang J, Ji L, et al (2022) Nüwa: Visual synthesis pre-training for neural visual world creation. In: European Conference on Computer Vision
2022
Later among the works it cites.
Xue H, Hang T, Zeng Y, et al (2022) Advancing high-resolution video-language representation with large-scale video transcriptions. In: Conference on Computer Vision and Pattern Recognition
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
Bertasius G, Wang H, Torresani L (2021) Is space-time attention all you need for video understanding? In: International Conference on Machine Learning
2021
Cited alongside, same era.
Dhariwal P, Nichol A (2021) Diffusion models beat GANs on image synthesis. In: Conference and Workshop on Neural Information Processing Systems
2021
Cited alongside, same era.
Ding M, Yang Z, Hong W, et al (2021) CogView: Mastering text-to-image generation via transformers. In: Conference and Workshop on Neural Information Processing Systems
2021
Cited alongside, same era.
Lee S, Chung J, Yu Y, et al (2021) ACAV100M: automatic curation of large-scale datasets for audio-visual video representation learning. In: International Conference on Computer Vision
2021
Cited alongside, same era.
Li Y, Zhang K, Cao J, et al (2021) Localvit: Bringing locality to vision transformers. arXiv
2021
Cited alongside, same era.
Menapace W, Lathuilière S, Tulyakov S, et al (2021) Playable video generation. In: Conference on Computer Vision and Pattern Recognition
2021
Cited alongside, same era.
Nichol A, Dhariwal P, Ramesh A, et al (2021) GLIDE: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv
2021
Cited alongside, same era.
2022
Later among the works it cites.
Zhou D, Wang W, Yan H, et al (2022) MagicVideo: Efficient video generation with latent diffusion models. arXiv
2022
Later among the works it cites.
An J, Zhang S, Yang H, et al (2023) Latent-Shift: Latent diffusion with temporal shift for efficient text-to-video generation. arXiv
2023
Closest in time.
Betker J, Goh G, Jing L, et al (2023) Improving image generation with better captions
2023
Closest in time.
Esser P, Chiu J, Atighehchian P, et al (2023) Structure and content-guided video synthesis with diffusion models. arXiv
2023
Closest in time.
Gu B, Fan H, Zhang L (2023) Two birds, one stone: A unified framework for joint learning of image and video style transfers. In: International Conference on Computer Vision
2023
Closest in time.
Kondratyuk D, Yu L, Gu X, et al (2023) Videopoet: A large language model for zero-shot video generation. arXiv
2023
Closest in time.
Li J, Li D, Savarese S, et al (2023) BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv
2023
Closest in time.
Luo Z, Chen D, Zhang Y, et al (2023) VideoFusion: Decomposed diffusion models for high-quality video generation. In: Conference on Computer Vision and Pattern Recognition
2023
Closest in time.
Ma Y, He Y, Cun X, et al (2023) Follow your pose: Pose-guided text-to-video generation using pose-free videos. arXiv
2023
Closest in time.
Ruan L, Ma Y, Yang H, et al (2023) MM-Diffusion: Learning multi-modal diffusion models for joint audio and video generation. In: Conference on Computer Vision and Pattern Recognition
2023
Closest in time.
Ruiz N, Li Y, Jampani V, et al (2023) DreamBooth: Fine tuning text-to-image diffusion models for subject-driven generation. In: Conference on Computer Vision and Pattern Recognition
2023
Closest in time.
Wang Y, Chen X, Ma X, et al (2023) LAVIE: high-quality video generation with cascaded latent diffusion models. arXiv
2023
Closest in time.
Wu JZ, Ge Y, Wang X, et al (2023) Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. In: International Conference on Computer Vision
2023
Closest in time.
Xu H, Ye Q, Yan M, et al (2023) mplug-2: A modularized multi-modal foundation model across text, image and video. In: International Conference on Machine Learning
2023
Closest in time.
Yang B, Gu S, Zhang B, et al (2023) Paint by example: Exemplar-based image editing with diffusion models. In: Conference on Computer Vision and Pattern Recognition
2023
Closest in time.
Yu S, Sohn K, Kim S, et al (2023) Video probabilistic diffusion models in projected latent space. In: Conference on Computer Vision and Pattern Recognition
2023
Closest in time.
Zhang DJ, Wu JZ, Liu J, et al (2023) Show-1: Marrying pixel and latent diffusion models for text-to-video generation. arXiv
2023
Closest in time.
Zhang L, Agrawala M (2023) Adding conditional control to text-to-image diffusion models. arXiv
2023
Closest in time.
Chen TS, Siarohin A, Menapace W, et al (2024) Panda-70m: Captioning 70m videos with multiple cross-modality teachers. In: Conference on Computer Vision and Pattern Recognition
2024
Closest in time.
Esser P, Kulal S, Blattmann A, et al (2024) Scaling rectified flow transformers for high-resolution image synthesis. arXiv
2024
Closest in time.
Guo Y, Yang C, Rao A, et al (2024) Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. International Conference on Learning Representations
2024
Closest in time.
Podell D, English Z, Lacey K, et al (2024) SDXL: improving latent diffusion models for high-resolution image synthesis. In: International Conference on Learning Representations
2024
Closest in time.