Fetching the paper…
Reading the bibliography…
Recent works have shown the capability of deep generative models to tackle general audio synthesis from a single label, producing a variety of impulsive, tonal, and environmental sounds.
Text-to-speech synthesis
Paul Taylor, · 2009
Earlier work this paper cites.
“A dataset and taxonomy for urban sound research,”
Justin Salamon, Christopher Jacoby, and Juan Pablo Bello, · 2014
Earlier work this paper cites.
“TUT database for acoustic scene classification and sound event detection,”
Annamaria Mesaros, Toni Heittola, and Tuomas Virtanen, · 2016
Earlier work this paper cites.
“PixelSNAIL: An improved autoregressive generative model,”
Xi Chen, Nikhil Mishra, Mostafa Rohaninejad, and Pieter Abbeel, · 2018
Earlier work this paper cites.
“A note on the inception score,”
Shane Barratt and Rishi Sharma, · 2018
Earlier work this paper cites.
“Acoustic scene generation with conditional SampleRNN,”
Qiuqiang Kong, Yong Xu, Turab Iqbal, Yin Cao, Wenwu Wang, and Mark D Plumbley, · 2019
Earlier work this paper cites.
“Generating diverse high-fidelity images with VQ-VAE-2,”
Ali Razavi, Aaron Van den Oord, and Oriol Vinyals, · 2019
Earlier work this paper cites.
“Generative modeling by estimating gradients of the data distribution,”
Yang Song and Stefano Ermon, · 2019
Earlier work this paper cites.
“Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms.,”
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi, · 2019
Earlier work this paper cites.
“Look, listen, and learn more: Design choices for deep audio embeddings,”
Jason Cramer, Ho-Hsiang Wu, Justin Salamon, and Juan Pablo Bello, · 2019
Earlier work this paper cites.
“Jukebox: A generative model for music,”
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever, · 2020
Cited alongside, same era.
“DrumGAN: Synthesis of drum sounds with timbral feature conditioning using generative adversarial networks,”
Javier Nistal, Stefan Lattner, and Gael Richard, · 2020
Cited alongside, same era.
“Hifi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,”
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, · 2020
Cited alongside, same era.
“DiffWave: A versatile diffusion model for audio synthesis,”
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro, · 2020
Cited alongside, same era.
“WaveGrad: Estimating gradients for waveform generation,”
Nanxin Chen, Yu Zhang, Heiga Zen, Ron J Weiss, Mohammad Norouzi, and William Chan, · 2020
Cited alongside, same era.
“Reliable fidelity and diversity metrics for generative models,”
“Conditional sound generation using neural discrete time-frequency representation learning,”
Xubo Liu, Turab Iqbal, Jinzheng Zhao, Qiushi Huang, Mark D Plumbley, and Wenwu Wang, · 2021
Later among the works it cites.
“Score-based generative modeling through stochastic differential equations,”
Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole, · 2021
Later among the works it cites.
“Adversarial score matching and improved sampling for image generation,”
Alexia Jolicoeur-Martineau, Rémi Piché-Taillefer, Rémi Tachet des Combes, and Ioannis Mitliagkas, · 2021
Later among the works it cites.
“On tuning consistent annealed sampling for denoising score matching,”
Joan Serrà, Santiago Pascual, and Jordi Pons, · 2021
Later among the works it cites.
“DiffSound: Discrete diffusion model for text-to-sound generation,”
Dongchao Yang, Jianwei Yu, Helin Wang, Wen Wang, Chao Weng, Yuexian Zou, and Dong Yu, · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Muhammad Ferjad Naeem, Seong Joon Oh, Youngjung Uh, Yunjey Choi, and Jaejun Yoo, · 2020
Cited alongside, same era.
“CRASH: Raw audio score-based generative modeling for controllable high-resolution drum sound synthesis,”
Simon Rouard and Gaëtan Hadjeres, · 2021
Cited alongside, same era.
“Neural synthesis of footsteps sound effects with generative adversarial networks,”
Marco Comunità, Huy Phan, and Joshua D Reiss, · 2021
Cited alongside, same era.
“Generating diverse realistic laughter for interactive art,”
M Mehdi Afsar, Eric Park, Étienne Paquette, Gauthier Gidel, Kory W Mathewson, and Eilif Muller, · 2021
Cited alongside, same era.
“AudioGen: Textually guided audio generation,”
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi, · 2022
Closest in time.
“Universal speech enhancement with score-based diffusion,”
Joan Serrà, Santiago Pascual, Jordi Pons, R Oguz Araz, and Davide Scaini, · 2022
Closest in time.
“Classifier-free diffusion guidance,”
Jonathan Ho and Tim Salimans, · 2022
Closest in time.
“Photorealistic text-to-image diffusion models with deep language understanding,”
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi, · 2022
Closest in time.