Fetching the paper…
Reading the bibliography…
Although being widely adopted for evaluating generated audio signals, the Fr\'echet Audio Distance (FAD) suffers from significant limitations, including reliance on Gaussian assumptions, sensitivity to sample size, and high computational complexity.
Hilbert space embeddings and metrics on probability measures
Bharath K Sriperumbudur, Arthur Gretton, Kenji Fukumizu, Bernhard Schölkopf, and Gert RG Lanckriet · 2010
Earlier work this paper cites.
Universality, characteristic kernels and rkhs embedding of measures
Bharath K Sriperumbudur, Kenji Fukumizu, and Gert RG Lanckriet · 2011
Earlier work this paper cites.
A kernel two-sample test
Arthur Gretton, Karsten M Borgwardt, Malte J Rasch, Bernhard Schölkopf, and Alexander Smola · 2012
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Cnn architectures for large-scale audio classification
Shawn Hershey, Sourish Chaudhuri, Daniel PW Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al · 2017
Earlier work this paper cites.
Fréchet audio distance: A metric for evaluating music enhancement algorithms
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi · 2018
Earlier work this paper cites.
Chris Donahue, Julian McAuley, and Miller Puckette · 2018
Earlier work this paper cites.
Mikołaj Bińkowski, Danica J Sutherland, Michael Arbel, and Arthur Gretton · 2018
Earlier work this paper cites.
Closed-form expressions for maximum mean discrepancy with applications to wasserstein auto-encoders
Raif M Rustamov · 2019
Earlier work this paper cites.
Audiocaps: Generating captions for audios in the wild
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim · 2019
Earlier work this paper cites.
Look, listen, and learn more: Design choices for deep audio embeddings
Aurora Linh Cramer, Ho-Hsiang Wu, Justin Salamon, and Juan Pablo Bello · 2019
Earlier work this paper cites.
Effectively unbiased fid and inception score and where to find them
Min Jin Chong and David Forsyth · 2020
Earlier work this paper cites.
Clotho: An audio captioning dataset
Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen · 2020
Earlier work this paper cites.
Panns: Large-scale pretrained audio neural networks for audio pattern recognition
Qiuqiang Kong, Yin Cao, Talha Iqbal, Yuxuan Wang, Zihao Yin, Wenwu Wang, and Mark D. Plumbley · 2020
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Earlier work this paper cites.
Comparing representations for audio synthesis using generative adversarial networks
Javier Nistal, Stefan Lattner, and Gaël Richard · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Cdpam: Contrastive learning for perceptual audio similarity
Pranay Manocha, Zeyu Jin, Richard Zhang, and Adam Finkelstein · 2021
Earlier work this paper cites.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed · 2021
Earlier work this paper cites.
Audiogen: Textually guided audio generation
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi · 2022
Earlier work this paper cites.
Efficient training of audio transformers with patchout
Khaled Koutini, Jan Schlüter, Hamid Eghbal-zadeh, and Gerhard Widmer · 2022
Earlier work this paper cites.
Wavlm: Large-scale self-supervised pre-training for full stack speech processing
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al · 2022
Earlier work this paper cites.
Audioldm: text-to-audio generation with latent diffusion models
Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, and Mark D Plumbley · 2023
Earlier work this paper cites.
Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models
Rongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren, Luping Liu, Mingze Li, Zhenhui Ye, Jinglin Liu, Xiang Yin, and Zhou Zhao · 2023
Earlier work this paper cites.
Text-to-audio generation using instruction guided latent diffusion model
Deepanway Ghosal, Navonil Majumder, Ambuj Mehrish, and Soujanya Poria · 2023
Earlier work this paper cites.
Accelerating diffusion-based text-to-audio generation with consistency distillation
Yatong Bai, Trung Dang, Dung Tran, Kazuhito Koishida, and Somayeh Sojoudi · 2023
Earlier work this paper cites.
Clipsonic: Text-to-audio synthesis with unlabeled videos and pretrained language-vision models
Hao-Wen Dong, Xiaoyu Liu, Jordi Pons, Gautam Bhattacharya, Santiago Pascual, Joan Serrà, Taylor Berg-Kirkpatrick, and Julian McAuley · 2023
Earlier work this paper cites.
I hear your true colors: Image guided audio generation
Roy Sheffer and Yossi Adi · 2023
Cited alongside, same era.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov · 2023
Cited alongside, same era.
Foley sound synthesis at the dcase 2023 challenge
Keunwoo Choi, Jaekwon Im, Laurie Heller, Brian Mcfee, Keisuke Imoto, Yuki Okamoto, Mathieu Lagrange, and Shinnosuke Takamichi · 2023
Cited alongside, same era.
Jless submission to dcase2023 task7: Foley sound synthesis using non-autoagressive generative model
Siwei Huang, Jisheng Bai, Yafei Jia, and Jianfeng Chen · 2023
Cited alongside, same era.
Hyu submission for the dcase 2023 task 7: Diffusion probabilistic model with adversarial training for foley sound synthesis
Won-Gook Choi and Joon-Hyuk Chang · 2023
Cited alongside, same era.
Audioldm 2: Learning holistic audio generation with self-supervised pretraining
Haohe Liu, Yi Yuan, Xubo Liu, Xinhao Mei, Qiuqiang Kong, Qiao Tian, Yuping Wang, Wenwu Wang, Yuxuan Wang, and Mark D Plumbley · 2024
Later among the works it cites.
Tango 2: Aligning diffusion-based text-to-audio generative models through direct preference optimization
Navonil Majumder, Chia-Yu Hung, Deepanway Ghosal, Wei-Ning Hsu, Rada Mihalcea, and Soujanya Poria · 2024
Later among the works it cites.
Improving text-to-audio models with synthetic captions
Zhifeng Kong, Sang-gil Lee, Deepanway Ghosal, Navonil Majumder, Ambuj Mehrish, Rafael Valle, Soujanya Poria, and Bryan Catanzaro · 2024
Later among the works it cites.
Auffusion: Leveraging the power of diffusion and large language models for text-to-audio generation
Jinlong Xue, Yayue Deng, Yingming Gao, and Ya Li · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fall-e: Gaudio foley synthesis system
Minsung Kang, Sangshin Oh, Hyeongi Moon, Kyungyun Lee, and Ben Sangbae Chon · 2023
Cited alongside, same era.
High-quality foley sound synthesis using monte carlo dropout
Chae-Woon Bang, Nam Kyun Kim, and Chanjun Chun · 2023
Cited alongside, same era.
Foley sound synthesis in waveform domain with diffusion model
Yoonjin Chung, Junwon Lee, and Juhan Nam · 2023
Cited alongside, same era.
Foley sound synthesis with audioldm for dcase2023 task 7
Shitong Fan, Qiaoxi Zhu, Feiyang Xiao, Haiyan Lan, Wenwu Wang, and Jian Guan1 · 2023
Cited alongside, same era.
Foley sound synthesis based on gan using contrastive learning without label information
Hae Chun Chung, Yuna Lee, and Jae Hoon Jung · 2023
Cited alongside, same era.
Dcase task-7: Stylegan2-based foley sound synthesis
Purnima Kamath, Tasnim Nishat Islam, Chitralekha Gupta, Lonce Wyse, and Suranga Nanayakkara · 2023
Cited alongside, same era.
Conditional foley sound synthesis with limited data: Two-stage data augmentation approach with stylegan2-ada
Kyungsu Kim, Jinwoo Lee, Hayoon Kim, and Kyogu Lee · 2023
Cited alongside, same era.
Zach Evans, Julian D Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons · 2024
Later among the works it cites.
Ezaudio: Enhancing text-to-audio generation with efficient diffusion transformer
Jiarui Hai, Yong Xu, Hao Zhang, Chenxing Li, Helin Wang, Mounya Elhilali, and Dong Yu · 2024
Later among the works it cites.
Syncfusion: Multimodal onset-synchronized video-to-audio foley synthesis
Marco Comunità, Riccardo F Gramaccioni, Emilian Postolache, Emanuele Rodolà, Danilo Comminiello, and Joshua D Reiss · 2024
Later among the works it cites.
Video-foley: Two-stage video-to-sound generation via temporal event condition for foley sound
Junwon Lee, Jaekwon Im, Dabin Kim, and Juhan Nam · 2024
Later among the works it cites.
Read, watch and scream! sound generation from text and video
Yujin Jeong, Yunji Kim, Sanghyuk Chun, and Jiyoung Lee · 2024
Later among the works it cites.
Frieren: Efficient video-to-audio generation with rectified flow matching
Yongqi Wang, Wenxiang Guo, Rongjie Huang, Jiawei Huang, Zehan Wang, Fuming You, Ruiqi Li, and Zhou Zhao · 2024
Later among the works it cites.
Temporally aligned audio for video with autoregression
Ilpo Viertola, Vladimir Iashin, and Esa Rahtu · 2024
Later among the works it cites.
Tell what you hear from what you see - video to audio generation through text
Xiulong Liu, Kun Su, and Eli Shlizerman · 2024
Later among the works it cites.
Video-guided foley sound generation with multimodal controls
Ziyang Chen, Prem Seetharaman, Bryan Russell, Oriol Nieto, David Bourgin, Andrew Owens, and Justin Salamon · 2024
Later among the works it cites.
Gotta hear them all: Sound source aware vision to audio generation
Wei Guo, Heng Wang, Weidong Cai, and Jianbo Ma · 2024
Later among the works it cites.
Vintage: Joint video and text conditioning for holistic audio generation
Saksham Singh Kushwaha and Yapeng Tian · 2024
Later among the works it cites.
Taming multimodal joint training for high-quality video-to-audio synthesis
Ho Kei Cheng, Masato Ishii, Akio Hayakawa, Takashi Shibuya, Alexander Schwing, and Yuki Mitsufuji · 2024
Later among the works it cites.
V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models
Heng Wang, Jianbo Ma, Santiago Pascual, Richard Cartwright, and Weidong Cai · 2024
Later among the works it cites.
Foleygen: Visually-guided audio generation
Xinhao Mei, Varun Nagaraja, Gael Le Lan, Zhaoheng Ni, Ernie Chang, Yangyang Shi, and Vikas Chandra · 2024
Later among the works it cites.
Tiva: Time-aligned video-to-audio generation
Xihua Wang, Yuyue Wang, Yihan Wu, Ruihua Song, Xu Tan, Zehua Chen, Hongteng Xu, and Guodong Sui · 2024
Later among the works it cites.
Diff-a-riff: Musical accompaniment co-creation via latent diffusion models
Javier Nistal, Marco Pasini, Cyran Aouameur, Maarten Grachten, and Stefan Lattner · 2024
Later among the works it cites.
Measuring audio prompt adherence with distribution-based embedding distances
Maarten Grachten · 2024
Later among the works it cites.
Versa: A versatile evaluation toolkit for speech, audio, and music
Jiatong Shi, Hye-jin Shim, Jinchuan Tian, Siddhant Arora, Haibin Wu, Darius Petermann, Jia Qi Yip, You Zhang, Yuxun Tang, Wangyou Zhang, et al · 2024
Later among the works it cites.
Presto! distilling steps and layers for accelerating music generation
Zachary Novack, Ge Zhu, Jonah Casebeer, Julian McAuley, Taylor Berg-Kirkpatrick, and Nicholas J Bryan · 2024
Later among the works it cites.
High-fidelity audio compression with improved rvqgan
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar · 2024
Later among the works it cites.
Challenge on sound scene synthesis: Evaluating text-to-audio generation
Junwon Lee, Modan Tailleur, Laurie M Heller, Keunwoo Choi, Mathieu Lagrange, Brian McFee, Keisuke Imoto, and Yuki Okamoto · 2024
Later among the works it cites.
Masked generative video-to-audio transformers with enhanced synchronicity
Santiago Pascual, Chunghsin Yeh, Ioannis Tsiamas, and Joan Serrà · 2025
Closest in time.
Audiomentations
Iver Jordal · 2025
Closest in time.
Sound scene synthesis at the dcase 2024 challenge
Mathieu Lagrange, Junwon Lee, Modan Tailleur, Laurie M. Heller, Keunwoo Choi, Brian McFee, Keisuke Imoto, and Yuki Okamoto · 2025
Closest in time.