Fetching the paper…
Reading the bibliography…
The diffusion models including Denoising Diffusion Probabilistic Models (DDPM) and score-based generative models have demonstrated excellent performance in speech synthesis tasks.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Tacotron: Towards End-to-End Speech Synthesis,”
Yuxuan Wang, R.J. Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J. Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, Quoc Le, Yannis Agiomyrgiannakis, Rob Clark, and Rif A. Saurous, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
Keith Ito and Linda Johnson, · 2017
Earlier work this paper cites.
“Neural ordinary differential equations,”
Ricky TQ Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud, · 2018
Earlier work this paper cites.
“Neural speech synthesis with transformer network,”
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu, · 2019
Earlier work this paper cites.
“Fastspeech: Fast, robust and controllable text to speech,”
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2019
Earlier work this paper cites.
“Denoising diffusion probabilistic models,”
Jonathan Ho, Ajay Jain, and Pieter Abbeel, · 2020
Earlier work this paper cites.
“Diffwave: A versatile diffusion model for audio synthesis,”
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro, · 2020
Earlier work this paper cites.
“Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,”
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, · 2020
Cited alongside, same era.
“Score-based generative modeling through stochastic differential equations,”
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole, · 2021
Cited alongside, same era.
“Diff-TTS: A Denoising Diffusion Model for Text-to-Speech,”
Myeonghun Jeong, Hyeongju Kim, Sung Jun Cheon, Byoung Jin Choi, and Nam Soo Kim, · 2021
Cited alongside, same era.
“Grad-tts: A diffusion probabilistic model for text-to-speech,”
Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, and Mikhail Kudinov, · 2021
Cited alongside, same era.
“Fastspeech 2: Fast and high-quality end-to-end text to speech,”
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2021
Cited alongside, same era.
“Flow matching for generative modeling,”
Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matthew Le, · 2022
Later among the works it cites.
“Flow straight and fast: Learning to generate and transfer data with rectified flow,”
Xingchao Liu, Chengyue Gong, and Qiang Liu, · 2022
Later among the works it cites.
“Interpretable Style Transfer for Text-to-Speech with ControlVAE and Diffusion Bridge,”
Wenhao Guan, Tao Li, Yishuang Li, Hukai Huang, Qingyang Hong, and Lin Li, · 2023
Closest in time.
“Comospeech: One-step speech and singing voice synthesis via consistency model,”
Zhen Ye, Wei Xue, Xu Tan, Jie Chen, Qifeng Liu, and Yike Guo, · 2023
Closest in time.
Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever, · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jinglin Liu, Chengxi Li, Yi Ren, Feiyang Chen, and Zhou Zhao, · 2022
Cited alongside, same era.
“Prodiff: Progressive fast diffusion model for high-quality text-to-speech,”
Rongjie Huang, Zhou Zhao, Huadai Liu, Jinglin Liu, Chenye Cui, and Yi Ren, · 2022
Cited alongside, same era.
“Diffgan-tts: High-fidelity and efficient text-to-speech with denoising diffusion gans,”
Songxiang Liu, Dan Su, and Dong Yu, · 2022
Cited alongside, same era.
“Resgrad: Residual denoising diffusion probabilistic models for text to speech,”
Zehua Chen, Yihan Wu, Yichong Leng, Jiawei Chen, Haohe Liu, Xu Tan, Yang Cui, Ke Wang, Lei He, Sheng Zhao, et al., · 2022
Cited alongside, same era.
Matthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer, Leda Sari, Rashel Moritz, Mary Williamson, Vimal Manohar, Yossi Adi, Jay Mahadeokar, et al., · 2023
Closest in time.
“Audioldm: Text-to-audio generation with latent diffusion models,”
Haohe Liu, Zehua Chen, Yiitan Yuan, Xinhao Mei, Xubo Liu, Danilo P. Mandic, Wenwu Wang, and MarkD . Plumbley, · 2023
Closest in time.
“DurIAN: Duration Informed Attention Network for Speech Synthesis,”
Chengzhu Yu, Heng Lu, Na Hu, Meng Yu, Chao Weng, Kun Xu, Peng Liu, Deyi Tuo, Shiyin Kang, Guangzhi Lei, Dan Su, and Dong Yu, · 2031
Closest in time.