Fetching the paper…
Reading the bibliography…
Denoising diffusion probabilistic models (DDPMs) are expressive generative models that have been used to solve a variety of speech synthesis problems.
Mel-cepstral distance measure for objective speech quality assessment
Kubichek, R · 1993
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Wang, Z., Bovik, A., Sheikh, H., and Simoncelli, E · 2003
Earlier work this paper cites.
Estimation of non-normalized statistical models by score matching
Hyvärinen, A · 2005
Earlier work this paper cites.
Springer Berlin Heidelberg, Berlin, Heidelberg, 2007
Dynamic Time Warping , pp. 69–84 · 2007
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Autoencoding beyond pixels using a learned similarity metric
Larsen, A. B. L., Sønderby, S. K., Larochelle, H., and Winther, O · 2016
Earlier work this paper cites.
Deconvolution and checkerboard artifacts
Odena, A., Dumoulin, V., and Olah, C · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Oord, A. v. d., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Deep voice 2: Multi-speaker neural text-to-speech
Gibiansky, A., Arik, S., Diamos, G., Miller, J., Peng, K., Ping, W., Raiman, J., and Zhou, Y · 2017
Earlier work this paper cites.
Least squares generative adversarial networks
Mao, X., Li, Q., Xie, H., Lau, R. Y., Wang, Z., and Paul Smolley, S · 2017
Earlier work this paper cites.
Char2wav: End-to-end speech synthesis
Sotelo, J., Mehri, S., Kumar, K., Santos, J. F., Kastner, K., Courville, A., and Bengio, Y · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Tacotron: Towards End-to-End Speech Synthesis
Wang, Y., Skerry-Ryan, R., Stanton, D., Wu, Y., Weiss, R. J., Jaitly, N., Yang, Z., Xiao, Y., Chen, Z., Bengio, S., Le, Q., Agiomyrgiannakis, Y., Clark, R., and Saurous, R. A · 2017
Earlier work this paper cites.
Searching for activation functions, 2018
Ramachandran, P., Zoph, B., and Le, Q. V · 2018
Cited alongside, same era.
A style-based generator architecture for generative adversarial networks
Karras, T., Laine, S., and Aila, T · 2019
Cited alongside, same era.
Melgan: Generative adversarial networks for conditional waveform synthesis
Kumar, K., Kumar, R., de Boissiere, T., Gestin, L., Teoh, W. Z., Sotelo, J., de Brébisson, A., Bengio, Y., and Courville, A. C · 2019
Cited alongside, same era.
Neural speech synthesis with transformer network
Li, N., Liu, S., Liu, Y., Zhao, S., and Liu, M · 2019
Cited alongside, same era.
Fastspeech: Fast, robust and controllable text to speech
Ren, Y., Ruan, Y., Tan, X., Qin, T., Zhao, S., Zhao, Z., and Liu, T.-Y · 2019
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Cited alongside, same era.
One tts alignment to rule them all
Badlani, R., Łancucki, A., Shih, K. J., Valle, R., Ping, W., and Catanzaro, B · 2021
Later among the works it cites.
Wavegrad: Estimating gradients for waveform generation
Chen, N., Zhang, Y., Zen, H., Weiss, R. J., Norouzi, M., and Chan, W · 2021
Later among the works it cites.
End-to-end adversarial text-to-speech
Donahue, J., Dieleman, S., Binkowski, M., Elsen, E., and Simonyan, K · 2021
Later among the works it cites.
Parallel tacotron: Non-autoregressive and controllable tts
Elias, I., Zen, H., Shen, J., Zhang, Y., Jia, Y., Weiss, R. J., and Wu, Y · 2021
Later among the works it cites.
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech
Kim, J., Kong, J., and Son, J · 2021
Later among the works it cites.
Diffwave: A versatile diffusion model for audio synthesis
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Glow-tts: A generative flow for text-to-speech via monotonic alignment search
Kim, J., Kim, S., Kong, J., and Yoon, S · 2020
Cited alongside, same era.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Kong, J., Kim, J., and Bae, J · 2020
Cited alongside, same era.
Flow-tts: A non-autoregressive network for text to speech based on flow
Miao, C., Liang, S., Chen, M., Ma, J., Wang, S., and Xiao, J · 2020
Cited alongside, same era.
Permutation invariant graph generation via score-based generative modeling
Niu, C., Song, Y., Song, J., Zhao, S., Grover, A., and Ermon, S · 2020
Cited alongside, same era.
Non-autoregressive neural text-to-speech
Peng, K., Ping, W., Song, Z., and Zhao, K · 2020
Cited alongside, same era.
VocGAN: A High-Fidelity Real-Time Vocoder with a Hierarchically-Nested Adversarial Network
Yang, J., Lee, J., Kim, Y., Cho, H.-Y., and Kim, I · 2020
Cited alongside, same era.
Kong, Z., Ping, W., Huang, J., Zhao, K., and Catanzaro, B · 2021
Later among the works it cites.
Bidirectional variational inference for non-autoregressive text-to-speech
Lee, Y., Shin, J., and Jung, K · 2021
Later among the works it cites.
Efficienttts: An efficient and high-quality text-to-speech architecture
Miao, C., Shuang, L., Liu, Z., Minchuan, C., Ma, J., Wang, S., and Xiao, J · 2021
Later among the works it cites.
Symbolic music generation with diffusion models
Mittal, G., Engel, J., Hawthorne, C., and Simon, I · 2021
Later among the works it cites.
Grad-tts: A diffusion probabilistic model for text-to-speech
Popov, V., Vovk, I., Gogoryan, V., Sadekova, T., and Kudinov, M · 2021
Later among the works it cites.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2021
Later among the works it cites.
Tackling the generative learning trilemma with denoising diffusion gans
Xiao, Z., Kreis, K., and Vahdat, A · 2021
Later among the works it cites.
GANSpeech: Adversarial Training for High-Fidelity Multi-Speaker Speech Synthesis
Yang, J., Bae, J.-S., Bak, T., Kim, Y.-I., and Cho, H.-Y · 2021
Later among the works it cites.