Fetching the paper…
Reading the bibliography…
We study Neural Foley, the automatic generation of high-quality sound effects synchronizing with videos, enabling an immersive audio-visual experience.
The Foley grail: The art of performing sound for film, games, and animation
Vanessa Theme Ament · 2014
Earlier work this paper cites.
Visually indicated sounds
Andrew Owens, Phillip Isola, Josh McDermott, Antonio Torralba, Edward H Adelson, and William T Freeman · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Visually indicated sound generation by perceptually optimized classification
Kan Chen, Chuanxi Zhang, Chen Fang, Zhaowen Wang, Trung Bui, and Ram Nevatia · 2018
Earlier work this paper cites.
Visual to sound: Generating natural sound for videos in the wild
Yipin Zhou, Zhaowen Wang, Chen Fang, Trung Bui, and Tamara L Berg · 2018
Earlier work this paper cites.
Vggsound: A large-scale audio-visual dataset
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman · 2020
Earlier work this paper cites.
Generating visually aligned sound from videos
Peihao Chen, Yang Zhang, Mingkui Tan, Hongdong Xiao, Deng Huang, and Chuang Gan · 2020
Earlier work this paper cites.
Audio-visual synchronisation in the wild
Honglie Chen, Weidi Xie, Triantafyllos Afouras, Arsha Nagrani, Andrea Vedaldi, and Andrew Zisserman · 2021
Earlier work this paper cites.
The benefit of temporally-strong labels in audio event classification
Shawn Hershey, Daniel PW Ellis, Eduardo Fonseca, Aren Jansen, Caroline Liu, R Channing Moore, and Manoj Plakal · 2021
Earlier work this paper cites.
Taming visually guided sound generation
Vladimir Iashin and Esa Rahtu · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Repetitive activity counting by sight and sound
Yunhua Zhang, Ling Shao, and Cees GM Snoek · 2021
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Cited alongside, same era.
Wav2clip: Learning robust audio representations from clip
Ho-Hsiang Wu, Prem Seetharaman, Kundan Kumar, and Juan Pablo Bello · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Clipsonic: Text-to-audio synthesis with unlabeled videos and pretrained language-vision models
Hao-Wen Dong, Xiaoyu Liu, Jordi Pons, Gautam Bhattacharya, Santiago Pascual, Joan Serrà, Taylor Berg-Kirkpatrick, and Julian McAuley · 2023
Cited alongside, same era.
Conditional generation of audio from video via foley analogies
Yuexi Du, Ziyang Chen, Justin Salamon, Bryan Russell, and Andrew Owens · 2023
Cited alongside, same era.
Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models
Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala · 2023
Later among the works it cites.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al · 2024
Closest in time.
Foley (filmmaking), 2024
Wikipedia contributors · 2024
Closest in time.
Freesound
Freesound Project · 2024
Closest in time.
Text-to-audio generation synchronized with videos
Shentong Mo, Jing Shi, and Yapeng Tian · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Text-to-audio generation using instruction-tuned llm and latent diffusion model
Deepanway Ghosal, Navonil Majumder, Ambuj Mehrish, and Soujanya Poria · 2023
Cited alongside, same era.
Imagebind: One embedding space to bind them all
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra · 2023
Cited alongside, same era.
Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models
Rongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren, Luping Liu, Mingze Li, Zhenhui Ye, Jinglin Liu, Xiang Yin, and Zhou Zhao · 2023
Cited alongside, same era.
Audioldm: text-to-audio generation with latent diffusion models
Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, and Mark D Plumbley · 2023
Cited alongside, same era.
Audioldm 2: Learning holistic audio generation with self-supervised pretraining
Haohe Liu, Qiao Tian, Yi Yuan, Xubo Liu, Xinhao Mei, Qiuqiang Kong, Yuping Wang, Wenwu Wang, Yuxuan Wang, and Mark D Plumbley · 2023
Cited alongside, same era.
Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models
Simian Luo, Chuanhao Yan, Chenxu Hu, and Hang Zhao · 2023
Cited alongside, same era.
I hear your true colors: Image guided audio generation
Roy Sheffer and Yossi Adi · 2023
Cited alongside, same era.
Closest in time.
Any-to-any generation via composable diffusion
Zineng Tang, Ziyi Yang, Chenguang Zhu, Michael Zeng, and Mohit Bansal · 2024
Closest in time.
V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models
Heng Wang, Jianbo Ma, Santiago Pascual, Richard Cartwright, and Weidong Cai · 2024
Closest in time.
Sonicvisionlm: Playing sound with vision language models
Zhifeng Xie, Shengye Yu, Mengtian Li, Qile He, Chaofeng Chen, and Yu-Gang Jiang · 2024
Closest in time.
Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners
Yazhou Xing, Yingqing He, Zeyue Tian, Xintao Wang, and Qifeng Chen · 2024
Closest in time.
Auffusion: Leveraging the power of diffusion and large language models for text-to-audio generation
Jinlong Xue, Yayue Deng, Yingming Gao, and Ya Li · 2024
Closest in time.
Audio-synchronized visual animation
Lin Zhang, Shentong Mo, Yijing Zhang, and Pedro Morgado · 2024
Closest in time.