Fetching the paper…
Reading the bibliography…
The field of text-to-audio generation has seen significant advancements, and yet the ability to finely control the acoustic characteristics of generated audio remains under-explored.
“Suppression of acoustic noise in speech using spectral subtraction,”
S. Boll, · 1979
Earlier work this paper cites.
An introduction to the psychology of hearing
Brian CJ Moore, · 2012
Earlier work this paper cites.
“Auto-encoding variational bayes,”
Diederik P Kingma, · 2013
Earlier work this paper cites.
“Generative adversarial nets,”
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Tacotron: Towards end-to-end speech synthesis,”
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al., · 2017
Earlier work this paper cites.
“Deep voice: Real-time neural text-to-speech,”
Sercan Ö. Arık, Mike Chrzanowski, Adam Coates, Gregory Diamos, Andrew Gibiansky, Yongguo Kang, Xian Li, John Miller, Andrew Ng, Jonathan Raiman, Shubho Sengupta, and Mohammad Shoeybi, · 2017
Earlier work this paper cites.
“Fr \ \backslash ’echet audio distance: A metric for evaluating music enhancement algorithms,”
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi, · 2018
Earlier work this paper cites.
“Hierarchical generative modeling for controllable speech synthesis,”
Wei-Ning Hsu, Yu Zhang, Ron J Weiss, Heiga Zen, Yonghui Wu, Yuxuan Wang, Yuan Cao, Ye Jia, Zhifeng Chen, Jonathan Shen, et al., · 2018
Earlier work this paper cites.
“Crepe: A convolutional representation for pitch estimation,”
Jong Wook Kim, Justin Salamon, Peter Li, and Juan Pablo Bello, · 2018
Earlier work this paper cites.
“Denoising diffusion probabilistic models,”
Jonathan Ho, Ajay Jain, and Pieter Abbeel, · 2020
Earlier work this paper cites.
“Diffwave: A versatile diffusion model for audio synthesis,”
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro, · 2021
Cited alongside, same era.
“Fsd50k: an open dataset of human-labeled sound events,”
Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, and Xavier Serra, · 2021
Cited alongside, same era.
“Audiogen: Textually guided audio generation,”
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi, · 2022
Cited alongside, same era.
“AudioLDM: Text-to-audio generation with latent diffusion models,”
Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, and Mark D Plumbley, · 2023
Cited alongside, same era.
“Neural codec language models are zero-shot text to speech synthesizers,”
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al., · 2023
Zach Evans, Julian D Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons, · 2024
Closest in time.
“Simple and controllable music generation,”
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez, · 2024
Closest in time.
“Compa: Addressing the gap in compositional reasoning in audio-language models,”
Sreyan Ghosh, Ashish Seth, Sonal Kumar, Utkarsh Tyagi, Chandra Kiran Reddy Evuru, Ramaneswaran S, S Sakshi, Oriol Nieto, Ramani Duraiswami, and Dinesh Manocha, · 2024
Closest in time.
“Autoregressive diffusion transformer for text-to-speech synthesis,”
Zhijun Liu, Shuai Wang, Sho Inoue, Qibing Bai, and Haizhou Li, · 2024
Closest in time.
“Music controlnet: Multiple time-varying controls for music generation,”
Shih-Lun Wu, Chris Donahue, Shinji Watanabe, and Nicholas J Bryan, · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“A demand-driven perspective on generative audio ai,”
Sangshin Oh, Minsung Kang, Hyeongi Moon, Keunwoo Choi, and Ben Sangbae Chon, · 2023
Cited alongside, same era.
“Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,”
Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov, · 2023
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al., · 2023
Cited alongside, same era.
“Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models,”
Rongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren, Luping Liu, Mingze Li, Zhenhui Ye, Jinglin Liu, Xiang Yin, and Zhou Zhao, · 2023
Cited alongside, same era.
“Scalable diffusion models with transformers,”
William Peebles and Saining Xie, · 2023
Cited alongside, same era.
“GAMA: A large audio-language model with advanced audio understanding and complex reasoning abilities,”
Sreyan Ghosh, Sonal Kumar, Ashish Seth, Chandra Kiran Reddy Evuru, Utkarsh Tyagi, S Sakshi, Oriol Nieto, Ramani Duraiswami, and Dinesh Manocha, · 2024
Closest in time.
“Scaling instruction-finetuned language models,”
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al., · 2024
Closest in time.
“High-fidelity audio compression with improved rvqgan,”
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar, · 2024
Closest in time.
“Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization,”
Navonil Majumder, Chia-Yu Hung, Deepanway Ghosal, Wei-Ning Hsu, Rada Mihalcea, and Soujanya Poria, · 2024
Closest in time.