Fetching the paper…
Reading the bibliography…
Denoising diffusion probabilistic models (DDPMs) have recently achieved leading performances in many generative tasks.
Mel-cepstral distance measure for objective speech quality assessment
Robert Kubichek. 1993 · 1993
Earlier work this paper cites.
Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs. In Proc. of ICASSP
Antony W Rix, John G Beerends, Michael P Hollier, and Andries P Hekstra. 2001 · 2001
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. 2004 · 2004
Earlier work this paper cites.
A short-time objective intelligibility measure for time-frequency weighted noisy speech. In Proc. of ICASSP
Cees H Taal, Richard C Hendriks, Richard Heusdens, and Jesper Jensen. 2010 · 2010
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning . PMLR, 2256–2265
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015 · 2015
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M Rush. 2016 · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. 2016 · 2016
Earlier work this paper cites.
Non-autoregressive neural machine translation
Jiatao Gu, James Bradbury, Caiming Xiong, Victor OK Li, and Richard Socher. 2017 · 2017
Earlier work this paper cites.
The lj speech dataset
Keith Ito. 2017 · 2017
Earlier work this paper cites.
Searching for activation functions
Prajit Ramachandran, Barret Zoph, and Quoc V Le. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Tacotron: Towards end-to-end speech synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al · 2017
Earlier work this paper cites.
Generative adversarial networks: An overview
Antonia Creswell, Tom White, Vincent Dumoulin, Kai Arulkumaran, Biswa Sengupta, and Anil A Bharath. 2018 · 2018
Earlier work this paper cites.
On GANs and GMMs. In Proc. of ICONIP
Eitan Richardson and Yair Weiss. 2018 · 2018
Earlier work this paper cites.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions. In Proc. of ICASSP . IEEE, 4779–4783
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al · 2018
Earlier work this paper cites.
Neural speech synthesis with transformer network. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 33. 6706–6713
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu. 2019 · 2019
Earlier work this paper cites.
Fastspeech: Fast, robust and controllable text to speech
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2019 · 2019
Earlier work this paper cites.
Token-level ensemble distillation for grapheme-to-phoneme conversion
Hao Sun, Xu Tan, Jun-Wei Gan, Hongzhi Liu, Sheng Zhao, Tao Qin, and Tie-Yan Liu. 2019 · 2019
Cited alongside, same era.
LibriTTS: A corpus derived from LibriSpeech for text-to-speech
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu. 2019 · 2019
Cited alongside, same era.
WaveGrad: Estimating Gradients for Waveform Generation. In Proc. of ICLR
Nanxin Chen, Yu Zhang, Heiga Zen, Ron J Weiss, Mohammad Norouzi, and William Chan. 2020 · 2020
Cited alongside, same era.
End-to-end adversarial text-to-speech
Jeff Donahue, Sander Dieleman, Mikołaj Bińkowski, Erich Elsen, and Karen Simonyan. 2020 · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Diffsinger: Diffusion acoustic model for singing voice synthesis
Jinglin Liu, Chengxi Li, Yi Ren, Feiyang Chen, Peng Liu, and Zhou Zhao. 2021 · 2021
Later among the works it cites.
Diffusion probabilistic models for 3d point cloud generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 2837–2845
Shitong Luo and Wei Hu. 2021 · 2021
Later among the works it cites.
Meta-stylespeech: Multi-speaker adaptive text-to-speech generation
Dongchan Min, Dong Bok Lee, Eunho Yang, and Sung Ju Hwang. 2021 · 2021
Later among the works it cites.
DiffSinger
MoonInTheRiver. 2021 · 2021
Later among the works it cites.
Grad-tts: A diffusion probabilistic model for text-to-speech. In International Conference on Machine Learning . PMLR, 8599–8608
Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, and Mikhail Kudinov. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Glow-tts: A generative flow for text-to-speech via monotonic alignment search
Jaehyeon Kim, Sungwon Kim, Jungil Kong, and Sungroh Yoon. 2020 · 2020
Cited alongside, same era.
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020a · 2020
Cited alongside, same era.
Fastspeech 2: Fast and high-quality end-to-end text to speech
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2020 · 2020
Cited alongside, same era.
Improved techniques for training score-based generative models
Yang Song and Stefano Ermon. 2020 · 2020
Cited alongside, same era.
Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram. In Proc. of ICASSP
Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim. 2020 · 2020
Cited alongside, same era.
Adaspeech: Adaptive text to speech for custom voice
Mingjian Chen, Xu Tan, Bohan Li, Yanqing Liu, Tao Qin, Sheng Zhao, and Tie-Yan Liu. 2021 · 2021
Cited alongside, same era.
EMOVIE: A Mandarin Emotion Speech Dataset with a Simple Emotional Text-to-Speech Model
Chenye Cui, Yi Ren, Jinglin Liu, Feiyang Chen, Rongjie Huang, Ming Lei, and Zhou Zhao. 2021 · 2021
Cited alongside, same era.
PortaSpeech: Portable and High-Quality Generative Text-to-Speech
Yi Ren, Jinglin Liu, and Zhou Zhao. 2021 · 2021
Later among the works it cites.
Noise estimation for generative diffusion models
Robin San-Roman, Eliya Nachmani, and Lior Wolf. 2021 · 2021
Later among the works it cites.
Tackling the Generative Learning Trilemma with Denoising Diffusion GANs
Zhisheng Xiao, Karsten Kreis, and Arash Vahdat. 2021 · 2021
Later among the works it cites.
GANSpeech: Adversarial Training for High-Fidelity Multi-Speaker Speech Synthesis
Jinhyeok Yang, Jae-Sung Bae, Taejun Bak, Youngik Kim, and Hoon-Young Cho. 2021 · 2021
Later among the works it cites.
GAN Vocoder: Multi-Resolution Discriminator Is All You Need
Jaeseong You, Dalhyun Kim, Gyuhyeon Nam, Geumbyeol Hwang, and Gyeongsu Chae. 2021 · 2021
Later among the works it cites.
FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis
Rongjie Huang, Max WY Lam, Jun Wang, Dan Su, Dong Yu, Yi Ren, and Zhou Zhao. 2022a · 2022
Closest in time.
GenerSpeech: Towards Style Transfer for Generalizable Out-Of-Domain Text-to-Speech Synthesis
Rongjie Huang, Yi Ren, Jinglin Liu, Chenye Cui, and Zhou Zhao. 2022b · 2022
Closest in time.
BDDM: Bilateral Denoising Diffusion Models for Fast and High-Quality Speech Synthesis. In Proc. of ICLR
Max WY Lam, Jun Wang, Dan Su, and Dong Yu. 2022 · 2022
Closest in time.
DiffGAN-TTS: High-Fidelity and Efficient Text-to-Speech with Denoising Diffusion GANs
Songxiang Liu, Dan Su, and Dong Yu. 2022 · 2022
Closest in time.
Revisiting Over-Smoothness in Text to Speech
Yi Ren, Xu Tan, Tao Qin, Zhou Zhao, and Tie-Yan Liu. 2022 · 2022
Closest in time.
Progressive distillation for fast sampling of diffusion models
Tim Salimans and Jonathan Ho. 2022 · 2022
Closest in time.