Fetching the paper…
Reading the bibliography…
Existing works have made strides in video generation, but the lack of sound effects (SFX) and background music (BGM) hinders a complete and immersive viewer experience.
Digital signal processing (3rd ed.): principles, algorithms, and applications
John G. Proakis and Dimitris G. Manolakis · 1996
Earlier work this paper cites.
Improved techniques for training gans
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Vggsound: A large-scale audio-visual dataset, 2020
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Earlier work this paper cites.
The benefit of temporally-strong labels in audio event classification, 2021
Shawn Hershey, Daniel P W Ellis, Eduardo Fonseca, Aren Jansen, Caroline Liu, R Channing Moore, and Manoj Plakal · 2021
Earlier work this paper cites.
Taming visually guided sound generation, 2021
Vladimir Iashin and Esa Rahtu · 2021
Earlier work this paper cites.
Audiogen: Textually guided audio generation
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi · 2022
Cited alongside, same era.
Simple and controllable music generation
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez · 2023
Cited alongside, same era.
Multimodal large language models: A survey, 2023
Jiayang Wu, Wensheng Gan, Zefeng Chen, Shicheng Wan, and Philip S. Yu · 2023
Cited alongside, same era.
Gemini: A family of highly capable multimodal models, 2023
Gemini Team Google · 2023
Cited alongside, same era.
Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models, 2023
Simian Luo, Chuanhao Yan, Chenxu Hu, and Hang Zhao · 2023
Cited alongside, same era.
I hear your true colors: Image guided audio generation, 2023
Roy Sheffer and Yossi Adi · 2023
Later among the works it cites.
Syncfusion: Multimodal onset-synchronized video-to-audio foley synthesis, 2023
Marco Comunità, Riccardo F. Gramaccioni, Emilian Postolache, Emanuele Rodolà, Danilo Comminiello, and Joshua D. Reiss · 2023
Later among the works it cites.
Foleygen: Visually-guided audio generation, 2023
Xinhao Mei, Varun Nagaraja, Gael Le Lan, Zhaoheng Ni, Ernie Chang, Yangyang Shi, and Vikas Chandra · 2023
Later among the works it cites.
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh · 2024
Closest in time.
Homepage, 2024
Pika · 2024
Closest in time.
Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Imagebind: One embedding space to bind them all, 2023
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra · 2023
Cited alongside, same era.
Conditional generation of audio from video via foley analogies, 2023
Yuexi Du, Ziyang Chen, Justin Salamon, Bryan Russell, and Andrew Owens · 2023
Cited alongside, same era.
Yazhou Xing, Yingqing He, Zeyue Tian, Xintao Wang, and Qifeng Chen · 2024
Closest in time.