Fetching the paper…
Reading the bibliography…
We propose the first joint audio-video generation framework that brings engaging watching and listening experiences simultaneously, towards high-quality realistic videos.
UCF101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling · 2014
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Quo vadis, action recognition? A new model and the kinetics dataset
João Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Deep cross-modal audio-visual generation
Lele Chen, Sudhanshu Srivastava, Zhiyao Duan, and Chenliang Xu · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
CMCGAN: A uniform framework for cross-modal visual-audio mutual generation
Wangli Hao, Zhaoxiang Zhang, and He Guan · 2018
Earlier work this paper cites.
Speech Commands: A dataset for limited-vocabulary speech recognition
Pete Warden · 2018
Earlier work this paper cites.
Visual to sound: Generating natural sound for videos in the wild
Yipin Zhou, Zhaowen Wang, Chen Fang, Trung Bui, and Tamara L. Berg · 2018
Earlier work this paper cites.
AIST dance video database: Multi-genre, multi-dancer, and multi-camera database for dance information processing
Shuhei Tsuchida, Satoru Fukayama, Masahiro Hamasaki, and Masataka Goto · 2019
Earlier work this paper cites.
Sound2Sight: Generating visual dynamics from sound and context
Moitreya Chatterjee and Anoop Cherian · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Learning texture transformer network for image super-resolution
Fuzhi Yang, Huan Yang, Jianlong Fu, Hongtao Lu, and Baining Guo · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol · 2021
Earlier work this paper cites.
Video background music generation with controllable music transformer
Shangzhe Di, Zeren Jiang, Si Liu, Zhaokai Wang, Leyan Zhu, Zexin He, Hongming Liu, and Shuicheng Yan · 2021
Earlier work this paper cites.
Domain-aware universal style transfer
Kibeom Hong, Seogkyu Jeon, Huan Yang, Jianlong Fu, and Hyeran Byun · 2021
Earlier work this paper cites.
Diffwave: A versatile diffusion model for audio synthesis
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro · 2021
Earlier work this paper cites.
AI Choreographer: music conditioned 3d dance generation with AIST++
Ruilong Li, Shan Yang, David A. Ross, and Angjoo Kanazawa · 2021
Cited alongside, same era.
More control for free! image synthesis with semantic diffusion guidance
Xihui Liu, Dong Huk Park, Samaneh Azadi, Gong Zhang, Arman Chopikyan, Yuxiao Hu, Humphrey Shi, Anna Rohrbach, and Trevor Darrell · 2021
Cited alongside, same era.
Improved denoising diffusion probabilistic models
Alexander Quinn Nichol and Prafulla Dhariwal · 2021
Cited alongside, same era.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2021
Cited alongside, same era.
Score-Based Generative Modeling through Stochastic Differential Equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole · 2021
Cited alongside, same era.
TTVFI: learning trajectory-aware transformer for video frame interpolation
Chengxu Liu, Huan Yang, Jianlong Fu, and Xueming Qian · 2022
Closest in time.
DPM-Solver: a fast ode solver for diffusion probabilistic model sampling in around 10 steps
Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu · 2022
Closest in time.
DPM-Solver++: fast solver for guided sampling of diffusion probabilistic models
Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu · 2022
Closest in time.
RePaint: Inpainting using denoising diffusion probabilistic models
Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool · 2022
Closest in time.
AI illustrator: Translating raw descriptions into images by prompt-based cross-modal generation
Yiyang Ma, Huan Yang, Bei Liu, Jianlong Fu, and Jiaying Liu · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hongwei Xue, Bei Liu, Huan Yang, Jianlong Fu, Houqiang Li, and Jiebo Luo · 2021
Cited alongside, same era.
Improving visual quality of image synthesis by A token-based generator with transformers
Yanhong Zeng, Huan Yang, Hongyang Chao, Jianbo Wang, and Jianlong Fu · 2021
Cited alongside, same era.
Learning conditional knowledge distillation for degraded-reference image quality assessment
Heliang Zheng, Huan Yang, Jianlong Fu, Zheng-Jun Zha, and Jiebo Luo · 2021
Cited alongside, same era.
Analytic-DPM: an analytic estimate of the optimal reverse variance in diffusion probabilistic models
Fan Bao, Chongxuan Li, Jun Zhu, and Bo Zhang · 2022
Cited alongside, same era.
Make-A-Scene: Scene-based text-to-image generation with human priors
Oran Gafni, Adam Polyak, Oron Ashual, Shelly Sheynin, Devi Parikh, and Yaniv Taigman · 2022
Cited alongside, same era.
Long video generation with time-agnostic VQGAN and time-sensitive transformer
Songwei Ge, Thomas Hayes, Harry Yang, Xi Yin, Guan Pang, David Jacobs, Jia-Bin Huang, and Devi Parikh · 2022
Cited alongside, same era.
AudioCLIP: Extending clip to image, text and audio
Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas Dengel · 2022
Cited alongside, same era.
Closest in time.
Learning spatiotemporal frequency-transformer for compressed video super-resolution
Zhongwei Qiu, Huan Yang, Jianlong Fu, and Dongmei Fu · 2022
Closest in time.
Hierarchical text-conditional image generation with CLIP latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Closest in time.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Closest in time.
Palette: Image-to-image diffusion models
Chitwan Saharia, William Chan, Huiwen Chang, Chris Lee, Jonathan Ho, Tim Salimans, David Fleet, and Mohammad Norouzi · 2022
Closest in time.
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al · 2022
Closest in time.
Image super-resolution via iterative refinement
Chitwan Saharia, Jonathan Ho, William Chan, Tim Salimans, David J Fleet, and Mohammad Norouzi · 2022
Closest in time.
Make-A-Video: Text-to-video generation without text-video data
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al · 2022
Closest in time.
Degradation-guided meta-restoration network for blind super-resolution
Fuzhi Yang, Huan Yang, Yanhong Zeng, Jianlong Fu, and Hongtao Lu · 2022
Closest in time.
Generating videos with dynamics-aware implicit generative adversarial networks
Sihyun Yu, Jihoon Tack, Sangwoo Mo, Hyunsu Kim, Junho Kim, Jung-Woo Ha, and Jinwoo Shin · 2022
Closest in time.
Discrete contrastive diffusion for cross-modal and conditional generation
Ye Zhu, Yu Wu, Kyle Olszewski, Jian Ren, Sergey Tulyakov, and Yan Yan · 2022
Closest in time.
AudioGen: Textually guided audio generation
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi · 2023
Closest in time.
Unified multi-modal latent diffusion for joint subject and text conditional image generation
Yiyang Ma, Huan Yang, Wenjing Wang, Jianlong Fu, and Jiaying Liu · 2023
Closest in time.
Accommodating audio modality in CLIP for multimodal processing
Ludan Ruan, Anwen Hu, Yuqing Song, Liang Zhang, Sipeng Zheng, and Qin Jin · 2023
Closest in time.