Fetching the paper…
Reading the bibliography…
We propose to synthesize high-quality and synchronized audio, given video and optional text conditions, using a novel multimodal joint training framework MMAudio.
A technique for the measurement of attitudes
Rensis Likert · 1932
Earlier work this paper cites.
A scale for the measurement of the psychological magnitude pitch
Stanley Smith Stevens, John Volkmann, and Edwin Broomell Newman · 1937
Earlier work this paper cites.
The fantastical gap between diegetic and nondiegetic, 2007
Robynn J Stilwell · 2007
Earlier work this paper cites.
Multisensory integration of drumming actions: musical expertise affects perceived audiovisual asynchrony
Karin Petrini, Sofia Dahl, Davide Rocchesso, Carl Haakon Waadeland, Federico Avanzini, Aina Puce, and Frank E Pollick · 2009
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Visually indicated sounds
Andrew Owens, Phillip Isola, Josh McDermott, Antonio Torralba, Edward H Adelson, and William T Freeman · 2016
Earlier work this paper cites.
Improved techniques for training gans
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen · 2016
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Self-normalizing neural networks
Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Fréchet audio distance: A metric for evaluating music enhancement algorithms
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi · 2018
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville · 2018
Earlier work this paper cites.
Audiocaps: Generating captions for audios in the wild
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Clotho: An audio captioning dataset
Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Panns: Large-scale pretrained audio neural networks for audio pattern recognition
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D Plumbley · 2020
Earlier work this paper cites.
Taming visually guided sound generation
Vladimir Iashin and Esa Rahtu · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans · 2022
Cited alongside, same era.
Efficient training of audio transformers with patchout
Khaled Koutini, Jan Schlüter, Hamid Eghbal-Zadeh, and Gerhard Widmer · 2022
Cited alongside, same era.
Flow matching for generative modeling
Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le · 2022
Cited alongside, same era.
Conditional generation of audio from video via foley analogies
Yuexi Du, Ziyang Chen, Justin Salamon, Bryan Russell, and Andrew Owens · 2023
Cited alongside, same era.
Text-to-audio generation using instruction tuned llm and latent diffusion model
Deepanway Ghosal, Navonil Majumder, Ambuj Mehrish, and Soujanya Poria · 2023
Cited alongside, same era.
Read, watch and scream! sound generation from text and video
Yujin Jeong, Yunji Kim, Sanghyuk Chun, and Jiyoung Lee · 2024
Closest in time.
Analyzing and improving the training dynamics of diffusion models
Tero Karras, Miika Aittala, Jaakko Lehtinen, Janne Hellsten, Timo Aila, and Samuli Laine · 2024
Closest in time.
A versatile diffusion transformer with mixture of noise levels for audiovisual generation
Gwanghyun Kim, Alonso Martinez, Yu-Chuan Su, Brendan Jou, José Lezama, Agrim Gupta, Lijun Yu, Lu Jiang, Aren Jansen, Jacob Walker, et al · 2024
Closest in time.
Conette: An efficient audio captioning system leveraging multiple datasets with task embedding
Etienne Labb, Thomas Pellegrini, and Julien Pinquier · 2024
Closest in time.
Audio dataset project
LAION-AI · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Imagebind: One embedding space to bind them all
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra · 2023
Cited alongside, same era.
Bigvgan: A universal neural vocoder with large-scale training
Sang-gil Lee, Wei Ping, Boris Ginsburg, Bryan Catanzaro, and Sungroh Yoon · 2023
Cited alongside, same era.
Audioldm: Text-to-audio generation with latent diffusion models
Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, and Mark D Plumbley · 2023
Cited alongside, same era.
Foleygen: Visually-guided audio generation
Xinhao Mei, Varun Nagaraja, Gael Le Lan, Zhaoheng Ni, Ernie Chang, Yangyang Shi, and Vikas Chandra · 2023
Cited alongside, same era.
Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation. 2023 ieee
Ludan Ruan, Y Ma, Huan Yang, Huiguo He, Bei Liu, Jianlong Fu, Nicholas Jing Yuan, Qin Jin, and Baining Guo · 2023
Cited alongside, same era.
Physics-driven diffusion models for impact sound synthesis from videos
Kun Su, Kaizhi Qian, Eli Shlizerman, Antonio Torralba, and Chuang Gan · 2023
Cited alongside, same era.
Improving and generalizing flow-based generative models with minibatch optimal transport
Alexander Tong, Nikolay Malkin, Guillaume Huguet, Yanlei Zhang, Jarrid Rector-Brooks, Kilian Fatras, Guy Wolf, and Yoshua Bengio · 2023
Cited alongside, same era.
Junwon Lee, Jaekwon Im, Dabin Kim, and Juhan Nam · 2024
Closest in time.
Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models
Simian Luo, Chuanhao Yan, Chenxu Hu, and Hang Zhao · 2024
Closest in time.
Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization
Navonil Majumder, Chia-Yu Hung, Deepanway Ghosal, Wei-Ning Hsu, Rada Mihalcea, and Soujanya Poria · 2024
Closest in time.
Towards temporally synchronized visually indicated sounds through scale-adapted positional embeddings
Xinhao Mei, Gael Le Lan, Haohe Liu, Zhaoheng Ni, Varun K Nagaraja, Anurag Kumar, Yangyang Shi, and Vikas Chandra · 2024
Closest in time.
Text-to-audio generation synchronized with videos
Shentong Mo, Jing Shi, and Yapeng Tian · 2024
Closest in time.
Masked generative video-to-audio transformers with enhanced synchronicity
Santiago Pascual, Chunghsin Yeh, Ioannis Tsiamas, and Joan Serrà · 2024
Closest in time.
Movie gen: A cast of media foundation models
Adam Polyak, Amit Zohar, Andrew Brown, Andros Tjandra, Animesh Sinha, Ann Lee, Apoorv Vyas, Bowen Shi, Chih-Yao Ma, Ching-Yao Chuang, et al · 2024
Closest in time.
Sta-v2a: Video-to-audio generation with semantic and temporal alignment
Yong Ren, Chenxing Li, Manjie Xu, Wei Liang, Yu Gu, Rilin Chen, and Dong Yu · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Closest in time.
Temporally aligned audio for video with autoregression
Ilpo Viertola, Vladimir Iashin, and Esa Rahtu · 2024
Closest in time.
Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners
Yazhou Xing, Yingqing He, Zeyue Tian, Xintao Wang, and Qifeng Chen · 2024
Closest in time.
Auffusion: Leveraging the power of diffusion and large language models for text-to-audio generation
Jinlong Xue, Yayue Deng, Yingming Gao, and Ya Li · 2024
Closest in time.
Draw an audio: Leveraging multi-instruction for video-to-audio synthesis
Qi Yang, Binjie Mao, Zili Wang, Xing Nie, Pengfei Gao, Ying Guo, Cheng Zhen, Pengfei Yan, and Shiming Xiang · 2024
Closest in time.
Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment
Bin Zhu, Bin Lin, Munan Ning, Yang Yan, Jiaxi Cui, HongFa Wang, Yatian Pang, Wenhao Jiang, Junwu Zhang, Zongwei Li, et al · 2024
Closest in time.